TII releases 1.6B-parameter Falcon-ASR: 20.92% average WER across six Arabic test sets, 22.73% on internal Emirati evaluation
Synopsis
The Technology Innovation Institute (TII) in Abu Dhabi released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic with particular attention to the Emirati dialect and also supporting English, French, Spanish and Portuguese; it reached 20.92% average word error rate across six Arabic test sets, better than the 23.17% best published result it cites, and 22.73% WER with 10.19% CER on an internal Emirati evaluation, the lowest among the systems compared, while also providing word-level timestamps.
Interpretation
Falcon-ASR reached 20.92% average word error rate and 8.79% average character error rate across six Arabic test sets, better than the 23.17% WER and 9.23% CER of the best published result in the same leaderboard snapshot, a gap of 2.25 percentage points. The result follows the Open Universal Arabic ASR Leaderboard's equal-weight averaging protocol using the leaderboard's pinned manifests, making systems of different sizes comparable under one measurement; the model has 1.6B parameters, fewer than the 2.0B to 7.0B systems compared. Evaluation follows the leaderboard protocol and its pinned manifests, and competitor figures are the published leaderboard averages checked on 30 September 2026; compared models include Audar-ASR-V1-Turbo, Cohere Transcribe Arabic and omniASR LLM 7B.
On TII's internal Emirati and Gulf speech evaluation, Falcon-ASR achieved 22.73% WER and 10.19% CER, the lowest among the systems compared, with WER 4.07 percentage points below the next best result, Qwen3-Omni. Public data already includes Emirati through the UAE subset of Casablanca; this internal evaluation complements that coverage with additional held-out recordings and human-validated transcripts, extending dialect assessment beyond the public UAE subset. The internal evaluation uses held-out recordings and human-validated transcripts and compares against Qwen3-Omni-30B-A3B-Instruct, Audar-ASR-V1-Turbo, Cohere Transcribe Arabic, Qwen3-ASR-1.7B-hf and Audar-ASR-V1-Flash on the same evaluation.
The same weights transcribe English, French, Spanish and Portuguese without a language flag, producing a transcript in the spoken language; on the seven public English test sets used by the Hugging Face Open ASR Leaderboard, mean WER was 5.74%. Multilingual capability is carried by a single set of model weights rather than switching models or prompts per language, and English results range from 1.75% on LibriSpeech clean to 11.86% on Earnings-22, spanning read, meeting and telephone settings. English results are reported per test set across the seven sets used by the public leaderboard; multilingual support is described as using the same weights without requiring a language flag.
Training covered background noise, overlapping speech, music, room reverberation and telephony effects, plus variations in speed and pitch, with the same treatment applied to Emirati recordings; the model also supports word-level timestamps linking each transcribed word to its position in the audio. These conditions target meetings, calls and other everyday recordings rather than only read or broadcast speech, and word-level timestamps add temporal alignment to the transcript. Training conditions and the timestamp capability are described directly by the releasing organization in the model write-up, without per-condition ablations or breakdown metrics.
Perspective
The results target speech transcription in Arabic, especially the Emirati dialect, plus English, French, Spanish and Portuguese, in settings such as meetings, calls and everyday recordings that include noise, overlapping speech, reverberation and telephony effects. The releasing organization offers a Hugging Face demo space for trying the model on one's own recordings, with API access and native applications planned. The model builds on the Falcon3-Audio work, whose architecture and training approach are described in the cited paper.
The recordings and transcripts of the internal Emirati evaluation are not released with the write-up, so outside readers cannot reproduce the 22.73% WER; competitor figures come from a leaderboard snapshot checked on 30 September 2026, and later leaderboard updates could change the ordering. Although training conditions list noise, reverberation and telephony effects, no per-condition ablation is given, so the individual contribution of each treatment is hard to judge. Accuracy metrics for word-level timestamps are not reported. In addition, this material is the releasing organization's own model write-up and does not include independent third-party reproduction.
