Service · Dialects

Arabic dialect data collection & annotation

Custom text and speech datasets across 25+ Arabic varieties, produced by native speakers of each one, labeled for dialect and task, and quality-reported variety by variety.

Arabic dialect data is text and speech in the Arabic people actually speak (Najdi, Hijazi, Egyptian, Levantine, Iraqi, Maghrebi and more), labeled so a model can be trained or tested on each variety. Bayanat Labs builds custom Arabic dialect datasets for teams shipping LLMs, speech recognition, chatbots and classifiers. Vetted native speakers of each variety produce the data; we label it for dialect and task and check it against per-variety gold sets. We cover 25+ Arabic varieties, host in-region by default, and start with a fixed-scope pilot, usually within two weeks.

Public dialect corpora are small, tweet- or phrasebook-based and often research-only. Our article on why dialect coverage sets your model's ceiling explains the gap.

25+Arabic varieties and dialects
Per varietygold sets and quality reporting
~2 weekstypical fixed-scope pilot

What we deliver

Dialect text collection

Chat, support requests, reviews and dialogue by native speakers of a named variety, with code-switching and Arabizi where your traffic has them.

Dialect speech collection

Scripted to spontaneous recordings with speaker quotas per variety. See Arabic speech data collection and transcription.

Dialect identification labels

Region, country or sub-variety labels at a declared granularity, multi-label where needed, plus level-of-dialectness scores.

Normalization & transliteration

Spelling fixed to a project lexicon, Arabizi converted to Arabic script, and dialect–MSA parallel pairs.

Task labels on dialect text

NER and spans, intent, sentiment and toxicity, labeled by annotators who speak the variety.

Per-variety evaluation sets

Held-out test sets sliced by dialect, not one blended score. See how to benchmark an Arabic LLM.

Which Arabic dialects? A variety-level coverage map

"Arabic" is not a sampling frame, and neither is "Gulf". Five-region schemes like LabelBench's (Egyptian, Gulf, Iraqi, Levantine, Maghrebi) suit benchmarks but hide differences users hear at once. Specify the level below:

GroupName these in a spec"What do you want?" (to a man)Labeling watch-out
MSA (reference)Formal written registerماذا تريد؟ mādhā turīd?Needs its own class
NajdiRiyadh, Qassim, Hailوش تبي؟ wesh tabi?Overlaps Gulf coast vocabulary; decide if "Saudi" is one label or several
HijaziJeddah, Makkah, Madinahإيش تبغى؟ ēsh tibgha?Urban b- prefix (إيش بتسوي؟ "what are you doing?") looks Egyptian or Levantine in short texts
Gulf (Khaleeji)Kuwaiti, Emirati, Qatari, Bahraini, Eastern Province, Omaniشتبي؟ sh-tabi? (Kuwaiti)k often becomes ch (چذي chidhi, "like this"), spelled with چ or ك
EgyptianCairene, Alexandrian, Saʿidiعايز إيه؟ ʿāyiz ēh?q is a glottal stop in Cairo, g in Upper Egypt; small LLMs over-predict Egyptian for other varieties (see LabelBench below)
LevantineSyrian, Lebanese, Jordanian, Palestinianشو بدّك؟ shū baddak?q shifts between city and rural speech within one country
IraqiBaghdadi, southern, Moslawiشتريد؟ shtrīd?"I said" is gilit in Baghdad, qeltu in Mosul
MaghrebiMoroccan Darija, Algerian, Tunisian, Libyan, Hassaniyaآش بغيتي؟ āsh bghīti? (Moroccan); شنوّة تحب؟ shnuwwa tḥebb? (Tunisian)n- marks "I" (نكتب nekteb, "I write"); heavy French code-switching and Arabizi
SudaneseKhartoum, central Sudanداير شنو؟ dāyir shinu?Absent from five-region schemes; decide early whether it gets its own label
YemeniSanaani, Taʿizzi-Adeni, Hadramiإيش تشتي؟ ēsh tishti? (Sanaani)Also absent from five-region schemes

Examples are common everyday forms in one of several accepted spellings. Dialects have no official orthography, so a gold set should list which variants count as correct.

Where public Arabic dialect datasets fall short

DatasetWhat it containsThe catch for production
MADAR (2018)2,000 sentences translated into 25 city dialects; 12,000 for five citiesTourist-phrasebook sentences averaging 6.5 words; licence covers internal research and evaluation only
QADI (2020)540,000 tweets from 18 countriesCountry labels come from users' self-declared profile descriptions, not from each tweet; 91.5% accurate on a manually checked random sample
NADI 20241,120 tweets, each labeled by three native speakers from each of ten countriesSmall; the provided training data was single-label, and the winning team scored 50.57 F1 on multi-label dialect ID

Single labels are the deeper flaw. When native speakers of seven Arabic dialects reviewed a dialect-ID model's mistakes, Keleg and Magdy (2023) found about 66% of validated errors were not true errors. Use public sets for baselines; for anything you ship, you need your domain, channel and variety mix, with rights you hold.

Arabic dialect data collection methods

Every producer is a screened native speaker of the target variety, and each item records their variety and region. We mix four methods:

  • Elicited: speakers render English or French prompts in their own dialect. MSA prompts pull output toward standard forms, which is why MADAR's authors avoided them.
  • Scenario-based: tasks like "dispute a card charge", for natural phrasing with planned domain coverage.
  • Conversational: chats or recordings between same-variety speakers, code-switching kept.
  • Scripted: entities, commands and numbers, for coverage you cannot leave to chance.

How we annotate dialect data

  • Dialect ID. Your granularity, plus MSA and mixed classes, with multiple labels allowed as in NADI 2024.
  • Normalization. The CODA* guidelines cover 28 city dialects; their authors found 27 attested spellings of one Egyptian word, مبيقولهاش (ma-bi-ʾul-hāsh, "he doesn't say it"). A project lexicon makes each word written, and scored, one way.
  • NER, intent and sentiment. Gulf ما عجبني ("I didn't like it") and Egyptian مش عاجبني ("I don't like it") both slip past a negation rule written for MSA لم يعجبني, so guidelines carry examples per variety (see how to annotate Arabic text).

How an Arabic dialect data project runs

  1. Scoping. Markets, varieties, task, channel, label granularity and schema.
  2. Guidelines and gold set. Per-variety examples and a lexicon; native linguists adjudicate gold items, and contributors train against them.
  3. Pilot. A fixed slice produced end to end, with a quality report by variety.
  4. QA and adjudication. Gold checks, double labeling and rework loops.
  5. Scale and delivery. Batches with producer metadata and an audit trail on every task: Vet → Train → Produce → QA → Deliver.

Per-variety quality reporting

Native linguists build a gold set for each variety. Gold items sit in every queue, and a share of items is labeled twice, independently, so agreement is measured per variety and per label. Recurring disagreements become guideline rules. Your report states the method, the producer mix and results by variety; we do not quote generic accuracy figures.

When to automate dialect labels and when to use native annotators

LabelBench, our reproducible audit, tested labelers on a balanced 500-item, five-region dialect task (Egyptian, Gulf, Iraqi, Levantine, Maghrebi) on the fixed held-out manifest, where a balanced guesser scores 20%:

LabelerExact-match accuracy
Qwen 2.5 3B, zero-shot22.4%
Fanar 1 9B (Arabic-focused), zero-shot47.8%
Character n-gram classifier trained on 500 disjoint labeled items (100 per region)69.8%

Qwen 2.5 3B called 75% of Gulf and 80% of Levantine examples Egyptian, and no dialect run cleared safe automation at a 95% quality floor. A simple classifier trained on 500 native, task-specific labeled examples outscored every LLM tested, which is why the labels themselves matter. Automate only the share of the queue that Safe Automation Coverage shows stays above your quality bar.

Security, consent and data residency

Consent wording and purpose are agreed at scoping; personal data is tagged or redacted. Hosting is in-region by default, with on-prem and private-cloud options. Access is least-privilege and audit-logged, and contributors work under NDA. See our PDPL and data residency checklist.

Use cases by industry

  • Telecom: chatbot and IVR intent models.
  • Finance and banking: complaint and sentiment classifiers.
  • Government: citizen-service assistants tested on Saudi varieties, not only MSA.
  • Healthcare: patient-message triage reviewed by licensed physicians.
  • Media and retail: moderation and search across dialects and Arabizi.

Public dataset, crowd platform or specialist vendor?

CriterionPublic datasetIn-house teamCrowd platformSpecialist vendor
Variety matchAs collectedYour staff's varietiesUsually self-reportedScreened natives per variety
RightsOften research-onlyYoursCheck the termsContractual
Dialect-level QAFixed at releaseDepends on bandwidthSpot checks or majority votePer-variety gold sets, adjudication
Best fitBaselines, researchSmall, sensitive scopesLarge, simple tasksMulti-dialect production data

Start with a pilot

Pick two or three varieties and one task, such as Najdi and Hijazi intent labels. The pilot has a fixed scope, gold-standard QA and a quality report by variety, usually within two weeks. Then scale: managed, embedded or enterprise. The same team runs broader Arabic data annotation and labeling. Tell us your markets.

Frequently asked questions

Is there a free Arabic dialect dataset I can use commercially?

A few, but read the licence first. MADAR, one of the most used, allows internal research and evaluation only, with no redistribution or modification. Tweet corpora add platform terms. Public sets suit baselines and experiments where the licence allows; for a shipped product, commission data whose rights you hold.

How fine-grained should Arabic dialect labels be?

As fine as the decisions your product makes. If one model serves all of Saudi Arabia, one Saudi label may do; if you route or localize by region, split Najdi, Hijazi and the rest. Add MSA and mixed classes, allow multiple labels, and pilot the scheme: a distinction native annotators cannot agree on is not worth paying to label.

Can machine translation produce Arabic dialect data?

Not reliably. Translating from English or MSA produces MSA-shaped sentences with foreign idiom underneath, and post-editing rarely removes all of it. Machine-drafted text can seed low-stakes experiments, but keep it out of evaluation sets. For why synthetic Arabic drifts toward the standard register, see model collapse speaks MSA.

Should an Arabic dialect dataset include Arabizi and code-switching?

If your users write that way, yes. Arabizi is Arabic typed in Latin script, with digits for some letters (3 for ع, 7 for ح), common in chat next to English or French. Decide in the spec whether to keep it as written, transliterate it, or both, and tag the language of each switched span.

Scope an Arabic dialect data pilot

Tell us your markets, varieties and task. We will propose a fixed-scope pilot with per-variety gold sets, gold-standard QA and a quality report, usually within two weeks.

Scope a dialect pilot