Skip to main content
Back to timeline
NVIDIA Technical BlogSource publication:

NVIDIA fine-tunes Nemotron 3.5 ASR on SADA and FLEURS, cutting Saudi Najdi and Hijazi WER from 55.05% to 29.96% while English also improves slightly

Synopsis

This NVIDIA tutorial adapts a Cache-Aware FastConformer-RNNT multilingual streaming ASR model to Najdi and Hijazi Saudi dialects using minimal curation of SADA 2022 (retaining 103,559 of 125,490 utterances, 133.7 hours, 82.5%), a weighted replay mix of 90% Saudi speech with 7% English and 3% Arabic FLEURS, and duration bucketing, training for 12,000 steps in about 4.5 hours on two GPUs, which lowers Najdi+Hijazi test WER from 55.05% to 29.96% and CER from 31.63% to 12.18%, full SADA WER from 58.84% to 35.61%, while FLEURS English WER falls from 11.04% to 10.42% and Arabic WER from 12.67% to 11.41%; it also compares encoder unfreezing depths (top 6 at 33.42%, top 8 at 32.32%, all 24 layers at 29.96%) and inference-time settings ([56,13] attention context gains 1.

AI-generated editorial illustration: Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages

Interpretation

Narrow-target fine-tuning on Saudi dialects with minimal curation substantially improves target-dialect recognition without sacrificing the dropped dialects or English. Compared with first fine-tuning on 11 dialects at once, where validation WER moved only from 49.5% to 46.7% and then plateaued, training only on Najdi and Hijazi lowered target test WER from 55.05% to 29.96% and full SADA WER from 58.84% to 35.61%, with FLEURS English and Arabic also slightly improving. WER/CER measured on independent test splits with the NeMo evaluation script; training ran 12,000 steps in about 4.5 hours on two GPUs; curation retained 103,559 of 125,490 utterances (133.7 hours, 82.5%).

Declaring replay proportions by weight (90% Saudi speech, 7% English, 3% Arabic FLEURS) is more reliable than concatenating replay files into a large manifest, and mitigates catastrophic forgetting. The text notes that a sliding-window shuffle may not reach rows appended to the end of a large manifest until late in training, so a concatenated replay set effectively does not exist for most of the run; switching to an explicit OmegaConf input_cfg with weights preserved English capability while the model specialized in Arabic dialect speech. FLEURS English WER fell from 11.04% to 10.42% and CER from 6.47% to 4.53%, while Arabic WER fell from 12.67% to 11.41%, indicating replay held existing languages while specialization proceeded.

Encoder unfreezing depth interacts with available data: at this experiment's 134 hours of target speech, unfreezing all 24 layers beat partial unfreezing. The text reports top 6 (WER 33.42%), top 8 (WER 32.32%), and all 24 layers (WER 29.96%), and records 230.4M trainable versus 407.6M frozen parameters in the top-eight recipe; partial freezing costs 2.4 points against the full fine-tune. A controlled comparison changing only the number of unfrozen layers under the same data and evaluation conditions; the authors state this is a finding about this data volume, not a general rule.

Inference-time settings trade latency for accuracy without retraining. Moving the attention context from the streaming default [56,3] to the widest [56,13] reduces WER by 1.31 absolute points at roughly 800 ms of extra buffering; MALSD beam-8 with [56,13] gains 2.71 points over greedy, and beam-4 reaches 28.81% WER at 0.59x greedy runtime. A decoding-configuration comparison table on the same checkpoint; the authors note MAES and NGPU-LM fusion were not evaluated, so the conclusion is limited to MALSD, and warn that strip_lang_tags=True is needed or the locale tag is scored as an insertion on every utterance.

Perspective

This pipeline is meant for teams that have enough labeled speech to specialize an ASR model but not enough to train one from scratch, covering dialect adaptation, domain-specific transcription, and deployments that must retain existing languages. The concrete numbers come from SADA 2022 Najdi and Hijazi plus FLEURS English and Arabic, with a Cache-Aware FastConformer-RNNT prompted multilingual streaming model (strip_lang_tags, target_lang: ar-AR), trained for 12,000 steps in about 4.5 hours on two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs. The authors frame the configuration changes as troubleshooting examples rather than recommended defaults, and note that weight decay, gradient clipping, mixed precision, and effective batch size in utterances were left at NeMo defaults and not tuned. The text also states that replay protects only what its data represents, that partial unfreezing needs re-tuning when the mix changes, and that the workflow does not generalize into evidence for every Arabic dialect or deployment environment. Moving to another language requires replacing the Arabic normalizer, validating Unicode normalization and encoding, testing the base tokenizer on names, numerals, borrowed words, and mixed-script sentences, and switching to character-, token-, or morpheme-level measures for languages without whitespace word boundaries.

Readers should note that all numbers come from a single dataset combination and a single model checkpoint, and the authors themselves stress that the configurations are troubleshooting records rather than recommended defaults, so corpora or hardware changes require re-validation. The replay proportions (7% English, 3% Arabic) and filtering thresholds (UTMOS >= 1.25, SIGMOS noise >= 1.5, SIGMOS overall >= 1.5) were set for this corpus distribution; the text notes a default UTMOS threshold of 3.0 would have rejected almost everything, indicating thresholds are highly corpus-sensitive. Decoding conclusions are limited to MALSD, with MAES and NGPU-LM fusion not evaluated. In addition, this is a tutorial-style article that does not report run-to-run variance or statistical significance, so differences on the order of 1 to 2 points (for example English WER from 11.04% to 10.42%) should be read as observations from this run. The text also mentions that NVIDIA Nemotron 3 Diarization extends the workflow to speaker-attributed transcription for up to 8 speakers, but architecture and benchmark details point to an external blog and are not developed here.

Sources