The UAE's Technology Innovation Institute (TII) released Falcon-Emirati-7B on Tuesday, October 6, 2026 — a language model that does something deceptively simple and surprisingly rare: it answers in Emirati dialect, not just Modern Standard Arabic. The official announcement, published on the institute's Hugging Face blog this morning, describes the project as teaching a model "the dialect, the culture, and the nuance" — and the benchmark gap it opens up over regional rivals Jais, ALLaM and Fanar is hard to ignore.
On Alyah, a native Emirati benchmark of 1,173 multiple-choice questions, the new model scored 84.83% — first among every model compared. More telling is dialect fidelity, judged with partial credit: Falcon-Emirati scored 0.52, while the nearest competitor, Saudi Arabia's ALLaM, managed 0.05 — roughly a ten-fold gap. Everyone else in the comparison, including Google's gemma-3-27b-it and the UAE's own Jais-2, fell below 0.05.

What TII actually built
Strip away the branding and the engineering story is straightforward. Falcon-Emirati-7B is a dialect-specialized layer of adaptation on top of the institute's existing Falcon-H1-Arabic family, which ships in 3B, 7B and 34B sizes. TII chose 7B deliberately: the 34B variant costs too much to train and serve for this purpose, while 3B "lacks cultural depth" in the team's own words.
Three technical details from the announcement matter for anyone evaluating the model:
- Hybrid architecture. Like the rest of the H1 line, each block runs Mamba (state-space) processing and Transformer attention in parallel, fusing their outputs — which is how the family supports very long contexts (128K, and up to 256K at family level).
- Three data channels: authentic Emirati-dialect web and forum crawls in native script (not transliterated text), MSA material on Emirati culture and heritage, and synthetic dialect data generated under strict glossary and style constraints.
- No established playbook existed. The team — Emirati researchers including Shaikha Alsuwaidi, Omar Saif Alkaabi and Maitha Alhammadi — reports that adapting models from MSA to dialect had no established recipe, so the data mix was found through ablations plus review by native speakers.
You can try it right now, free, on the institute's Falcon Chat portal by selecting Falcon-Emirati-7B from the model list — no subscription, no technical setup.
The numbers behind the ranking
Four separate evaluations, all published by TII, point the same direction:
- Alyah accuracy (1,173 native questions, benchmark published openly): 84.83%, ahead of every Arabic and multilingual model tested.
- UAE culture scenarios (283 dialogues from the ArabCulture-Dialogue suite): 85.57%, ahead of ALLaM at 83.39%, Jais-2 at 73.79% and Qatar's Fanar-2 at 71.50%.
- Head-to-head judging against Fanar-2-27B — a model four times its size: Falcon-Emirati-7B wins every category, from poetry (0.88 to 0.12) to religious and social sensitivity (0.80 to 0.20).
- Abstention: Fanar declines to answer 26.2% of questions; every other model stays under 5%.
One observation from the team is worth quoting indirectly: competitors often "know" the right answer but default to MSA even when prompted in Emirati — they lose on dialect, not knowledge. The judge model for the fidelity evaluations was Gemini 3.7 Flash, and TII acknowledges its weakest category is greetings and daily expressions, precisely where MSA and Emirati overlap most.


Why this matters if you make content
Gulf audiences respond to dialect the way English audiences respond to a natural, conversational voice — it signals belonging rather than broadcasting. Until now, creators producing Arabic content had to choose between general models that flatten everything into stiff MSA and manual rewriting that eats the day. A model that holds Emirati voice opens several concrete workflows:
- Scripts and captions for Reels and Shorts targeting UAE viewers, written natively rather than translated.
- Community management: replies that match how people actually talk in the comments.
- Localization QA: a second opinion on whether your "Gulf Arabic" copy sounds Emirati, Saudi or generic.
- Brand safety in tone: the model's strong showing on cultural sensitivity scenarios means fewer accidental missteps in religious or social contexts.
A practical stack for regional content is emerging: large general models for research and structure — we covered OpenAI's latest moves with GPT-6.1 Sol and Google's recent Gemini free-tier changes — plus specialized models like this one for voice. If you want that workflow unified — Arabic writing, images and multi-platform scheduling in one place — ARWriter's auto-writer is built exactly for it.
Quick comparison: the Arabic model field today
| Model | Maker | Dialect fidelity* | UAE culture | Note |
|---|---|---|---|---|
| Falcon-Emirati-7B | TII (UAE) | 0.52 | 85.57% | Clear leader in dialect and culture |
| ALLaM-7B | SDAIA (Saudi Arabia) | 0.05 | 83.39% | Close on culture, weak on dialect |
| Jais-2-8B | G42 (UAE) | 0.02 | 73.79% | Only wins daily greetings |
| Fanar-2-27B | QF (Qatar) | ≈0.00 | 71.50% | Abstains on 26% of questions |
* LLM-judged with partial credit on Alyah — higher is better. All figures from TII's official blog post.
Honest limitations
- One size only. There is no Emirati 34B variant; heavyweight reasoning still needs a larger model.
- No stated license. The announcement does not specify weight licensing or commercial API availability — budget for uncertainty if you plan product integration, and watch TII's channels for updates.
- TII's own caveats: the model can err on rare expressions and localized references; dialect nuance is "subjective even among native speakers"; sensitive deployments deserve domain-specific evaluation first.
- Self-reported benchmarks. The evaluations were run by TII with a Google model as judge. The institute explicitly invites the Emirati community to stress-test and report issues.
- Emirati, not pan-Gulf. The base family covers Gulf, Levantine, Egyptian and Maghrebi Arabic, but the specialization announced today is for the Emirati dialect specifically.
How to stress-test it in ten minutes
Laboratory benchmarks are one thing; your workload is the real test. A quick protocol any creator can run on the Falcon Chat portal:
- Tone test: ask for a Reel intro about a Dubai café in Emirati dialect, then run the identical brief through a large general model. Compare which output a Gulf viewer would find natural without explanation.
- Culture test: ask for appropriate compliment phrases for a formal Emirati occasion, or the difference between two everyday expressions — this is where the heritage-data training should show.
- Consistency test: repeat the same brief three times with different phrasing. Genuine dialect competence survives rephrasing; synthetic "dialect-flavored" MSA collapses back to formal Arabic on at least one attempt.
Note what happens on long answers, and how it handles a question written in Egyptian or Levantine dialect — the base family trained across those varieties, so behavior there tells you how it fits a pan-Arab workflow. And if you want a like-for-like baseline, the underlying Falcon-H1-Arabic family is documented separately with all three sizes.
Frequently asked questions
What is Falcon-Emirati-7B?
A 7-billion-parameter language model from the UAE's Technology Innovation Institute, specialized in the Emirati dialect and cultural context. It is built on the Falcon-H1-Arabic family and is free to try on the official Falcon Chat portal.
How good is it at Emirati dialect compared to other Arabic models?
In TII's evaluations it scored 0.52 on judged dialect fidelity versus 0.05 for the nearest competitor, and it won every category of a head-to-head comparison against the much larger Fanar-2-27B.
Can I use Falcon-Emirati-7B commercially?
Chat access is open to everyone, but the announcement does not state a license for the weights or commercial API terms. Check TII's official channels before building a product on it.
Does it work for other Gulf dialects?
The underlying Falcon-H1-Arabic family was pretrained across Gulf, Levantine, Egyptian and Maghrebi Arabic, but the dialect specialization, benchmarks and cultural tuning announced today are Emirati-specific.
Where can I try it?
On the official chat portal at chat.falconllm.tii.ae — select Falcon-Emirati-7B from the model picker. A quick dialect test takes minutes and tells you more than any chart.
Bottom line: Falcon-Emirati-7B marks a shift in Arabic AI from models that understand Arabic to models that understand a specific audience — and the first beneficiaries are creators and brands speaking to the Gulf in its own voice. Watch TII's official newsroom for any follow-up on licensing or larger variants. For a complete Arabic content workflow — writing, visuals and scheduling in one platform — take a look at ARWriter.