Not established from the available sources.
NVIDIA · MODEL RELEASE TRACKER
NVIDIA Nemotron 3.5 ASR
NVIDIA Nemotron 3.5 ASR is an automatic speech recognition model supporting multilingual streaming transcription across 40 language-locales. The model is adaptable via fine-tuning (the posted workflow uses 133.7 hours of Saudi Arabic data) and can trade off accuracy, latency, and compute by changing encoder updates, lookahead frames, and decoding settings.CURRENT SNAPSHOT3/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Lookahead frames and decodingUsing a larger attention context of 13 lookahead frames and beam-8 MALSD decoding lowered WER by 2.71 absolute points without retraining, at the cost of approximately 800 ms additional latency (suitable for batch transcription workloads).
ModalityAutomatic speech recognition (multilingual streaming transcription).
Saudi Arabic (Najdi/Hijazi) WER after fine-tuningFine-tuning on 133.7 hours of Najdi and Hijazi speech reduced word error rate from 55.05% to 29.96% on the target test split.
English WER change after adaptationEnglish performance improved from 11.04% to 10.42% WER following the adaptation described.
Not established from the available sources.
VERIFIABLE FACTS
Every value stays attached to a source and date.
MODALITIES · ModalityDEVELOPER CLAIM
Automatic speech recognition (multilingual streaming transcription).
CAPABILITIES · Supported language-localesDEVELOPER CLAIM
Supports multilingual streaming transcription across 40 language-locales, including transcription-ready Arabic.
BENCHMARKS · Saudi Arabic (Najdi/Hijazi) WER after fine-tuningDEVELOPER CLAIM
Fine-tuning on 133.7 hours of Najdi and Hijazi speech reduced word error rate from 55.05% to 29.96% on the target test split.
BENCHMARKS · English WER change after adaptationDEVELOPER CLAIM
English performance improved from 11.04% to 10.42% WER following the adaptation described.
CONTEXT WINDOW · Lookahead frames and decodingDEVELOPER CLAIM
Using a larger attention context of 13 lookahead frames and beam-8 MALSD decoding lowered WER by 2.71 absolute points without retraining, at the cost of approximately 800 ms additional latency (suitable for batch transcription workloads).
CAPABILITIES · Encoder updates and trainable parametersDEVELOPER CLAIM
Updating all 24 encoder layers achieved the lowest error rates but required 230.4 million trainable parameters; freezing layers is offered to reduce compute when data or memory are limited. The workflow uses techniques such as weighted replay mixing, duration-based bucketing, and partial encoder unfreezing.
LIMITATIONS · Dialect and deployment limitationsDEVELOPER CLAIM
Although the model performs across many languages, deployment-specific dialects and local recording conditions (for example Najdi and Hijazi) can cause poor performance unless targeted fine-tuning is applied; broad pretraining does not guarantee strong results on underrepresented regional dialects.
WHAT CHANGED