RELEASE · MODELS · #862
NVIDIA releases Nemotron 3 Diarization — open-weight 100M model for real-time multi-speaker diarization
NVIDIA published Nemotron 3 Diarization, an open-weight, 100M-parameter speaker-diarization model that ranks #1 on Voice Arena's Diarization-Bench (14.72% DER). The model supports up to eight anonymous speaker channels, handles overlapping speech in streaming and offline modes, and uses arrival-order speaker caching (AOSC) and a FIFO context buffer; training included public and licensed data, with David AI data reducing compound DER by 0.77 points.
KEY POINTS
- NVIDIA published Nemotron 3 Diarization, an open-weight, 100M-parameter speaker-diarization model that ranks #1 on Voice Arena's Diarization-Bench (14.72% DER).
- The model supports up to eight anonymous speaker channels, handles overlapping speech in streaming and offline modes, and uses arrival-order speaker caching (AOSC) and a FIFO context buffer; training included public and licensed data, with David AI data reducing compound DER by 0.77 points.
- A top-ranked, open-weight streaming diarization model that scales to eight speakers can materially improve speaker-attributed transcripts and downstream meeting/podcast analytics and ASR pipelines.
WHY IT MATTERS
A top-ranked, open-weight streaming diarization model that scales to eight speakers can materially improve speaker-attributed transcripts and downstream meeting/podcast analytics and ASR pipelines.