Tech Meridian← ALL MODELS
PROMY MERIDIAN

NVIDIA · MODEL RELEASE TRACKER

Nemotron 3 Diarization

Nemotron 3 Diarization is a Nvidia-released diarization model (~100M parameters) that identifies which speaker is talking in recordings or live audio, supports up to eight speakers, detects overlapping speech, and whose weights are freely available.

CURRENT SNAPSHOT4/5 DIMENSIONS WITH DATA

The dimensions that change the decision.

PRICING

Not established from the available sources.

CONTEXT WINDOW
Audio buffer settingsAudio buffer configurable at four levels ranging from 30.4 down to 0.32 seconds; shorter buffers generally reduce accuracy.
MODALITIES
Primary taskSpeaker diarization (who spoke when) — audio diarization
Input modalitiesWorks with both recordings and live audio (streaming).
BENCHMARKS
Benchmark performanceRanks #1 on Voice Arena's Diarization-Bench with a 14.72% Diarization Error Rate (DER).
Benchmark performanceRanked first on the Diarization-Bench (VoiceArena) with a 14.72% error rate; the next best system had 19.3%. The benchmark counts overlapping speech and small misalignments as errors.
AVAILABILITY
Weights availabilityModel weights are freely available.

RELEASE TIMELINE

Published, source-backed release events only.

1 SOURCES · IMPORTANCE 7.0

Nvidia releases Nemotron 3 Diarization, a 100M-parameter real-time speaker diarization model

Nvidia released Nemotron 3 Diarization, a ~100 million-parameter model whose weights are freely available. The model can identify up to eight speakers (including overlapping speech), works on live and recorded audio with configurable buffer sizes, and achieves a 14.72% error rate on VoiceArena's Diarization-Bench, outperforming the prior best system and reducing error versus Streaming Sortformer by about 41% on certain tests.

→
1 SOURCES · IMPORTANCE 7.0

NVIDIA releases Nemotron 3 Diarization — open-weight 100M model for real-time multi-speaker diarization

NVIDIA published Nemotron 3 Diarization, an open-weight, 100M-parameter speaker-diarization model that ranks #1 on Voice Arena's Diarization-Bench (14.72% DER). The model supports up to eight anonymous speaker channels, handles overlapping speech in streaming and offline modes, and uses arrival-order speaker caching (AOSC) and a FIFO context buffer; training included public and licensed data, with David AI data reducing compound DER by 0.77 points.

→

VERIFIABLE FACTS

Every value stays attached to a source and date.

WHAT CHANGED

Stored passport versions, without reconstructed history.

Passport updated12 facts
CONTEXT WINDOW · Audio buffer settings+ Audio buffer configurable at four levels ranging from 30.4 down to 0.32 seconds; shorter buffers generally reduce accuracy.
BENCHMARKS · Benchmark performance+ Ranked first on the Diarization-Bench (VoiceArena) with a 14.72% error rate; the next best system had 19.3%. The benchmark counts overlapping speech and small misalignments as errors.
MODALITIES · Input modalities+ Works with both recordings and live audio (streaming).
CAPABILITIES · Maximum number of speakers+ Can tell apart up to eight speakers.
CAPABILITIES · Overlapping speech detection+ Can detect when multiple people talk at the same time (overlapping speech).
CAPABILITIES · Parameter count100M parameters→About 100 million parameters.
RELEASE · Release announcement+ Reportedly released by Nvidia (article published 2026-09-27).
AVAILABILITY · Weights− Open-weight (weights released / open)
AVAILABILITY · Weights availability+ Model weights are freely available.
Passport created6 facts