Arabic dialect data collection & annotation
Custom text and speech datasets across 25+ Arabic varieties, produced by native speakers of each one, labeled for dialect and task, and quality-reported variety by variety.
Arabic dialect data is text and speech in the Arabic people actually speak (Najdi, Hijazi, Egyptian, Levantine, Iraqi, Maghrebi and more), labeled so a model can be trained or tested on each variety. Bayanat Labs builds custom Arabic dialect datasets for teams shipping LLMs, speech recognition, chatbots and classifiers. Vetted native speakers of each variety produce the data; we label it for dialect and task and check it against per-variety gold sets. We cover 25+ Arabic varieties, host in-region by default, and start with a fixed-scope pilot, usually within two weeks.
Public dialect corpora are small, tweet- or phrasebook-based and often research-only. Our article on why dialect coverage sets your model's ceiling explains the gap.
What we deliver
Dialect text collection
Chat, support requests, reviews and dialogue by native speakers of a named variety, with code-switching and Arabizi where your traffic has them.
Dialect speech collection
Scripted to spontaneous recordings with speaker quotas per variety. See Arabic speech data collection and transcription.
Dialect identification labels
Region, country or sub-variety labels at a declared granularity, multi-label where needed, plus level-of-dialectness scores.
Normalization & transliteration
Spelling fixed to a project lexicon, Arabizi converted to Arabic script, and dialect–MSA parallel pairs.
Task labels on dialect text
NER and spans, intent, sentiment and toxicity, labeled by annotators who speak the variety.
Per-variety evaluation sets
Held-out test sets sliced by dialect, not one blended score. See how to benchmark an Arabic LLM.
Which Arabic dialects? A variety-level coverage map
"Arabic" is not a sampling frame, and neither is "Gulf". Five-region schemes like LabelBench's (Egyptian, Gulf, Iraqi, Levantine, Maghrebi) suit benchmarks but hide differences users hear at once. Specify the level below:
| Group | Name these in a spec | "What do you want?" (to a man) | Labeling watch-out |
|---|---|---|---|
| MSA (reference) | Formal written register | ماذا تريد؟ mādhā turīd? | Needs its own class |
| Najdi | Riyadh, Qassim, Hail | وش تبي؟ wesh tabi? | Overlaps Gulf coast vocabulary; decide if "Saudi" is one label or several |
| Hijazi | Jeddah, Makkah, Madinah | إيش تبغى؟ ēsh tibgha? | Urban b- prefix (إيش بتسوي؟ "what are you doing?") looks Egyptian or Levantine in short texts |
| Gulf (Khaleeji) | Kuwaiti, Emirati, Qatari, Bahraini, Eastern Province, Omani | شتبي؟ sh-tabi? (Kuwaiti) | k often becomes ch (چذي chidhi, "like this"), spelled with چ or ك |
| Egyptian | Cairene, Alexandrian, Saʿidi | عايز إيه؟ ʿāyiz ēh? | q is a glottal stop in Cairo, g in Upper Egypt; small LLMs over-predict Egyptian for other varieties (see LabelBench below) |
| Levantine | Syrian, Lebanese, Jordanian, Palestinian | شو بدّك؟ shū baddak? | q shifts between city and rural speech within one country |
| Iraqi | Baghdadi, southern, Moslawi | شتريد؟ shtrīd? | "I said" is gilit in Baghdad, qeltu in Mosul |
| Maghrebi | Moroccan Darija, Algerian, Tunisian, Libyan, Hassaniya | آش بغيتي؟ āsh bghīti? (Moroccan); شنوّة تحب؟ shnuwwa tḥebb? (Tunisian) | n- marks "I" (نكتب nekteb, "I write"); heavy French code-switching and Arabizi |
| Sudanese | Khartoum, central Sudan | داير شنو؟ dāyir shinu? | Absent from five-region schemes; decide early whether it gets its own label |
| Yemeni | Sanaani, Taʿizzi-Adeni, Hadrami | إيش تشتي؟ ēsh tishti? (Sanaani) | Also absent from five-region schemes |
Examples are common everyday forms in one of several accepted spellings. Dialects have no official orthography, so a gold set should list which variants count as correct.
Where public Arabic dialect datasets fall short
| Dataset | What it contains | The catch for production |
|---|---|---|
| MADAR (2018) | 2,000 sentences translated into 25 city dialects; 12,000 for five cities | Tourist-phrasebook sentences averaging 6.5 words; licence covers internal research and evaluation only |
| QADI (2020) | 540,000 tweets from 18 countries | Country labels come from users' self-declared profile descriptions, not from each tweet; 91.5% accurate on a manually checked random sample |
| NADI 2024 | 1,120 tweets, each labeled by three native speakers from each of ten countries | Small; the provided training data was single-label, and the winning team scored 50.57 F1 on multi-label dialect ID |
Single labels are the deeper flaw. When native speakers of seven Arabic dialects reviewed a dialect-ID model's mistakes, Keleg and Magdy (2023) found about 66% of validated errors were not true errors. Use public sets for baselines; for anything you ship, you need your domain, channel and variety mix, with rights you hold.
Arabic dialect data collection methods
Every producer is a screened native speaker of the target variety, and each item records their variety and region. We mix four methods:
- Elicited: speakers render English or French prompts in their own dialect. MSA prompts pull output toward standard forms, which is why MADAR's authors avoided them.
- Scenario-based: tasks like "dispute a card charge", for natural phrasing with planned domain coverage.
- Conversational: chats or recordings between same-variety speakers, code-switching kept.
- Scripted: entities, commands and numbers, for coverage you cannot leave to chance.
How we annotate dialect data
- Dialect ID. Your granularity, plus MSA and mixed classes, with multiple labels allowed as in NADI 2024.
- Normalization. The CODA* guidelines cover 28 city dialects; their authors found 27 attested spellings of one Egyptian word, مبيقولهاش (ma-bi-ʾul-hāsh, "he doesn't say it"). A project lexicon makes each word written, and scored, one way.
- NER, intent and sentiment. Gulf ما عجبني ("I didn't like it") and Egyptian مش عاجبني ("I don't like it") both slip past a negation rule written for MSA لم يعجبني, so guidelines carry examples per variety (see how to annotate Arabic text).
How an Arabic dialect data project runs
- Scoping. Markets, varieties, task, channel, label granularity and schema.
- Guidelines and gold set. Per-variety examples and a lexicon; native linguists adjudicate gold items, and contributors train against them.
- Pilot. A fixed slice produced end to end, with a quality report by variety.
- QA and adjudication. Gold checks, double labeling and rework loops.
- Scale and delivery. Batches with producer metadata and an audit trail on every task: Vet → Train → Produce → QA → Deliver.
Per-variety quality reporting
Native linguists build a gold set for each variety. Gold items sit in every queue, and a share of items is labeled twice, independently, so agreement is measured per variety and per label. Recurring disagreements become guideline rules. Your report states the method, the producer mix and results by variety; we do not quote generic accuracy figures.
When to automate dialect labels and when to use native annotators
LabelBench, our reproducible audit, tested labelers on a balanced 500-item, five-region dialect task (Egyptian, Gulf, Iraqi, Levantine, Maghrebi) on the fixed held-out manifest, where a balanced guesser scores 20%:
| Labeler | Exact-match accuracy |
|---|---|
| Qwen 2.5 3B, zero-shot | 22.4% |
| Fanar 1 9B (Arabic-focused), zero-shot | 47.8% |
| Character n-gram classifier trained on 500 disjoint labeled items (100 per region) | 69.8% |
Qwen 2.5 3B called 75% of Gulf and 80% of Levantine examples Egyptian, and no dialect run cleared safe automation at a 95% quality floor. A simple classifier trained on 500 native, task-specific labeled examples outscored every LLM tested, which is why the labels themselves matter. Automate only the share of the queue that Safe Automation Coverage shows stays above your quality bar.
Security, consent and data residency
Consent wording and purpose are agreed at scoping; personal data is tagged or redacted. Hosting is in-region by default, with on-prem and private-cloud options. Access is least-privilege and audit-logged, and contributors work under NDA. See our PDPL and data residency checklist.
Use cases by industry
- Telecom: chatbot and IVR intent models.
- Finance and banking: complaint and sentiment classifiers.
- Government: citizen-service assistants tested on Saudi varieties, not only MSA.
- Healthcare: patient-message triage reviewed by licensed physicians.
- Media and retail: moderation and search across dialects and Arabizi.
Public dataset, crowd platform or specialist vendor?
| Criterion | Public dataset | In-house team | Crowd platform | Specialist vendor |
|---|---|---|---|---|
| Variety match | As collected | Your staff's varieties | Usually self-reported | Screened natives per variety |
| Rights | Often research-only | Yours | Check the terms | Contractual |
| Dialect-level QA | Fixed at release | Depends on bandwidth | Spot checks or majority vote | Per-variety gold sets, adjudication |
| Best fit | Baselines, research | Small, sensitive scopes | Large, simple tasks | Multi-dialect production data |
Start with a pilot
Pick two or three varieties and one task, such as Najdi and Hijazi intent labels. The pilot has a fixed scope, gold-standard QA and a quality report by variety, usually within two weeks. Then scale: managed, embedded or enterprise. The same team runs broader Arabic data annotation and labeling. Tell us your markets.
Frequently asked questions
Is there a free Arabic dialect dataset I can use commercially?
A few, but read the licence first. MADAR, one of the most used, allows internal research and evaluation only, with no redistribution or modification. Tweet corpora add platform terms. Public sets suit baselines and experiments where the licence allows; for a shipped product, commission data whose rights you hold.
How fine-grained should Arabic dialect labels be?
As fine as the decisions your product makes. If one model serves all of Saudi Arabia, one Saudi label may do; if you route or localize by region, split Najdi, Hijazi and the rest. Add MSA and mixed classes, allow multiple labels, and pilot the scheme: a distinction native annotators cannot agree on is not worth paying to label.
Can machine translation produce Arabic dialect data?
Not reliably. Translating from English or MSA produces MSA-shaped sentences with foreign idiom underneath, and post-editing rarely removes all of it. Machine-drafted text can seed low-stakes experiments, but keep it out of evaluation sets. For why synthetic Arabic drifts toward the standard register, see model collapse speaks MSA.
Should an Arabic dialect dataset include Arabizi and code-switching?
If your users write that way, yes. Arabizi is Arabic typed in Latin script, with digits for some letters (3 for ع, 7 for ح), common in chat next to English or French. Decide in the spec whether to keep it as written, transliterate it, or both, and tag the language of each switched span.