Falcon-Emirati-7B: Dialect-specialized LLM for Emirati Arabic

On October 6, 2026 the authors released Falcon-Emirati-7B, a 7B-parameter model adapted from Falcon-H1-Arabic to understand and generate Emirati Arabic more naturally. The team describes the effort as a targeted dialect adaptation rather than a full new base model.

Architecture and rationale

Falcon-Emirati-7B is built on the Falcon-H1 hybrid architecture used in Falcon-H1-Arabic, which combines State Space Models (Mamba) and Transformer attention in parallel inside each block, with outputs fused before projection. The Falcon-H1-Arabic family spans 3B, 7B, and 34B variants and supports context windows up to 128K and 256K tokens. The authors chose the 7B variant as a balance between capacity for dialect nuance and practical training and inference cost, noting the 34B option would raise costs while the 3B lacked sufficient headroom.

Data and adaptation strategy

The reported training pipeline used three complementary data sources: (1) curated authentic Emirati-dialect web content written natively in the dialect; (2) MSA material focused on Emirati culture, history and social norms to provide context; and (3) synthetic Emirati-dialect data generated under strict glossaries and style rules to fill topical gaps. The authors describe substantial experimentation—ablations on data mixes, training stages, and supervision strategies—to find the most effective adaptation recipe and to avoid overfitting to synthetic patterns.

For evaluation they combined native-speaker manual review with automated benchmarks. The team used Alyah, a 1,173-sample multiple-choice benchmark for Emirati-dialect capability, and also ran open-ended generation scored by an LLM judge (Gemini 3.7 Flash).

Results reported by the authors include an Alyah accuracy of 84.83%, which they state outperformed the other compared Arabic and multilingual models. In LLM-judged open-ended tests, Falcon-Emirati-7B led on correctness and showed a dialect-fidelity partial-credit score of 0.52 versus much lower scores for ALLaM-7B-Instruct-preview, gemma-3-27b-it, Jais-2-8B-Chat and Fanar-2-27B-Instruct. The team also reports Falcon-Emirati-7B scored 85.57% on the UAE portion of ArabCulture-Dialogue.

The authors provide category-level analyses showing the model’s advantage is strongest on poetry, figurative language and language-and-dialect tasks, and note that competing models often default to Modern Standard Arabic even when Emirati is expected. The paper and benchmark details are presented alongside the model release.


Original source: Hugging Face Blog

Leave a Comment