RELEASE · MODELS · #1087
Nvidia releases Nemotron 3 Diarization, a 100M-parameter real-time speaker diarization model
Nvidia released Nemotron 3 Diarization, a ~100 million-parameter model whose weights are freely available. The model can identify up to eight speakers (including overlapping speech), works on live and recorded audio with configurable buffer sizes, and achieves a 14.72% error rate on VoiceArena's Diarization-Bench, outperforming the prior best system and reducing error versus Streaming Sortformer by about 41% on certain tests.
KEY POINTS
- Nvidia released Nemotron 3 Diarization, a ~100 million-parameter model whose weights are freely available.
- The model can identify up to eight speakers (including overlapping speech), works on live and recorded audio with configurable buffer sizes, and achieves a 14.72% error rate on VoiceArena's Diarization-Bench, outperforming the prior best system and reducing error versus Streaming Sortformer by about 41% on certain tests.
- Open weights and state-of-the-art diarization accuracy make it easier to add real-time, multi‑speaker labeling to speech systems and research.
WHY IT MATTERS
Open weights and state-of-the-art diarization accuracy make it easier to add real-time, multi‑speaker labeling to speech systems and research.