Nvidia releases Nemotron 3 Diarization, a 100M-parameter real-time speaker diarization model
Nvidia released Nemotron 3 Diarization, a ~100 million-parameter model whose weights are freely available. The model can identify up to eight speakers (including overlapping speech), works on live and recorded audio with configurable buffer sizes, and achieves a 14.72% error rate on VoiceArena's Diarization-Bench, outperforming the prior best system and reducing error versus Streaming Sortformer by about 41% on certain tests.