Falcon-ASR: TII’s 1.6B Arabic ASR Targeting the Emirati Dialect

Falcon-ASR is a 1.6 billion‑parameter automatic speech recognition model developed at the Technology Innovation Institute (TII) in Abu Dhabi, with particular focus on the Emirati dialect. The authors report an average word error rate (WER) of 20.92% across six Arabic test sets and note a lowest-in-comparison result on an internal Emirati evaluation.

Benchmark results and comparisons

The evaluation follows the Open Universal Arabic ASR Leaderboard protocol, measuring equal-weight average WER across six public Arabic test sets and reporting character error rate (CER). The authors state Falcon-ASR’s average WER of 20.92% improves on the best published leaderboard snapshot figure of 23.17% (competitor averages checked on 30 September 2026).

On a held-out internal Emirati and Gulf speech evaluation, Falcon-ASR recorded a 22.73% WER and 10.19% CER. The authors report these were the lowest WER and CER among the systems they compared, with Falcon-ASR’s Emirati WER 4.07 percentage points below Qwen3-Omni, the next best result in that comparison.

Training, languages and features

Training data included Emirati, Modern Standard Arabic (MSA), other Gulf and Arabic dialects, plus English. The developers applied augmentation for background noise, overlapping speech, music, room reverberation, telephony effects and variations in speed and pitch; the same processing was applied to Emirati recordings to expose the model to common real‑world conditions.

The model produces word‑level timestamps that link each transcribed word to its audio position. The authors also report English transcription performance: a mean WER of 5.74% over seven public English test sets used by the Hugging Face Open ASR Leaderboard. Falcon-ASR uses the same weights to support Arabic, English, French, Spanish and Portuguese without requiring an explicit language flag; outputs are returned in the language actually spoken.

Falcon-ASR is built on the Falcon3-Audio work; the architecture and training approach are described in the paper “Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data.” A demo is available on the Hugging Face Demo Space, and the authors say API access and native applications are planned. The team acknowledges contributions from the Falcon-Emirati developers and thanks Mikhail Lubinets for compute‑infrastructure support.


Original source: Hugging Face Blog

Leave a Comment