NEWS · MODELS · #1345
NVIDIA fine-tunes Nemotron 3.5 ASR on Saudi Najdi and Hijazi dialects
Using the NeMo framework, NVIDIA fine-tuned Nemotron 3.5 ASR on 133.7 hours of Najdi and Hijazi speech, cutting WER on the target test split from 55.05% to 29.96% and slightly improving English WER from 11.04% to 10.42% without degrading other Arabic dialects. The workflow uses minimal curation, weighted replay with FLEURS, duration-based bucketing, and partial encoder unfreezing; switching to 13 lookahead frames plus beam-8 MALSD decoding reduced WER by an additional 2.71 points at ~800 ms extra latency, and Nemotron 3 Diarization adds speaker-attributed transcription for up to eight speakers.
KEY POINTS
- Using the NeMo framework, NVIDIA fine-tuned Nemotron 3.5 ASR on 133.7 hours of Najdi and Hijazi speech, cutting WER on the target test split from 55.05% to 29.96% and slightly improving English WER from 11.04% to 10.42% without degrading other Arabic dialects.
- The workflow uses minimal curation, weighted replay with FLEURS, duration-based bucketing, and partial encoder unfreezing; switching to 13 lookahead frames plus beam-8 MALSD decoding reduced WER by an additional 2.71 points at ~800 ms extra latency, and Nemotron 3 Diarization adds speaker-attributed transcription for up to eight speakers.
- Shows a practical recipe and trade-offs for adapting a large multilingual ASR to underrepresented regional dialects, yielding substantial WER reductions while preserving other-language performance.
WHY IT MATTERS
Shows a practical recipe and trade-offs for adapting a large multilingual ASR to underrepresented regional dialects, yielding substantial WER reductions while preserving other-language performance.