Not established from the available sources.
NVIDIA · MODEL RELEASE TRACKER
Nemotron 3 Diarization
Nemotron 3 Diarization is a Nvidia-released diarization model (~100M parameters) that identifies which speaker is talking in recordings or live audio, supports up to eight speakers, detects overlapping speech, and whose weights are freely available.CURRENT SNAPSHOT4/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
RELEASE TIMELINE
Published, source-backed release events only.
Nvidia releases Nemotron 3 Diarization, a 100M-parameter real-time speaker diarization model
Nvidia released Nemotron 3 Diarization, a ~100 million-parameter model whose weights are freely available. The model can identify up to eight speakers (including overlapping speech), works on live and recorded audio with configurable buffer sizes, and achieves a 14.72% error rate on VoiceArena's Diarization-Bench, outperforming the prior best system and reducing error versus Streaming Sortformer by about 41% on certain tests.
NVIDIA releases Nemotron 3 Diarization — open-weight 100M model for real-time multi-speaker diarization
NVIDIA published Nemotron 3 Diarization, an open-weight, 100M-parameter speaker-diarization model that ranks #1 on Voice Arena's Diarization-Bench (14.72% DER). The model supports up to eight anonymous speaker channels, handles overlapping speech in streaming and offline modes, and uses arrival-order speaker caching (AOSC) and a FIFO context buffer; training included public and licensed data, with David AI data reducing compound DER by 0.77 points.
VERIFIABLE FACTS
Every value stays attached to a source and date.
Speaker diarization (who spoke when) — audio diarization
About 100 million parameters.
Supports up to eight speakers in live and recorded conversations
Handles overlapping speech; supports chunked processing for flexible recording lengths; customizable streaming latency; orders output speakers by first appearance (Sortformer approach).
Model weights are freely available.
Ranks #1 on Voice Arena's Diarization-Bench with a 14.72% Diarization Error Rate (DER).
Reportedly released by Nvidia (article published 2026-09-27).
Can tell apart up to eight speakers.
Can detect when multiple people talk at the same time (overlapping speech).
Works with both recordings and live audio (streaming).
Audio buffer configurable at four levels ranging from 30.4 down to 0.32 seconds; shorter buffers generally reduce accuracy.
Ranked first on the Diarization-Bench (VoiceArena) with a 14.72% error rate; the next best system had 19.3%. The benchmark counts overlapping speech and small misalignments as errors.
WHAT CHANGED