Tech Meridian← ALL MODELS
PROMY MERIDIAN

NVIDIA · MODEL RELEASE TRACKER

NVIDIA Nemotron 3.5 ASR

NVIDIA Nemotron 3.5 ASR is an automatic speech recognition model supporting multilingual streaming transcription across 40 language-locales. The model is adaptable via fine-tuning (the posted workflow uses 133.7 hours of Saudi Arabic data) and can trade off accuracy, latency, and compute by changing encoder updates, lookahead frames, and decoding settings.

CURRENT SNAPSHOT3/5 DIMENSIONS WITH DATA

The dimensions that change the decision.

PRICING

Not established from the available sources.

CONTEXT WINDOW
Lookahead frames and decodingUsing a larger attention context of 13 lookahead frames and beam-8 MALSD decoding lowered WER by 2.71 absolute points without retraining, at the cost of approximately 800 ms additional latency (suitable for batch transcription workloads).
MODALITIES
ModalityAutomatic speech recognition (multilingual streaming transcription).
BENCHMARKS
Saudi Arabic (Najdi/Hijazi) WER after fine-tuningFine-tuning on 133.7 hours of Najdi and Hijazi speech reduced word error rate from 55.05% to 29.96% on the target test split.
English WER change after adaptationEnglish performance improved from 11.04% to 10.42% WER following the adaptation described.
AVAILABILITY

Not established from the available sources.

VERIFIABLE FACTS

Every value stays attached to a source and date.

CAPABILITIES · Encoder updates and trainable parametersDEVELOPER CLAIM

Updating all 24 encoder layers achieved the lowest error rates but required 230.4 million trainable parameters; freezing layers is offered to reduce compute when data or memory are limited. The workflow uses techniques such as weighted replay mixing, duration-based bucketing, and partial encoder unfreezing.

WHAT CHANGED

Stored passport versions, without reconstructed history.

Passport created7 facts