Arabic Speech Recognition Datasets Compared: Dialects, Hours & Licences
Most public Arabic speech is broadcast news, read prompts or YouTube, and much of it is licensed for research only. Here is what each major corpus actually contains, what you can use commercially, and where the gaps sit for production ASR.
Hours, dialects and licences below are taken from each dataset's own card, catalogue page or paper, checked in September 2026. Cards change and new versions ship; check the licence before commercial use, and treat this page as a map, not legal advice.
The short answer: the largest public Arabic speech recognition datasets are QASR (about 2,000 hours) and MGB-2 (about 1,200 hours) of Al Jazeera broadcasts, followed by MASC (1,000 hours of YouTube audio) and SADA (about 667 hours of Saudi television). Common Voice, FLEURS, Casablanca, Mixat and a few telephone corpora from the LDC fill smaller niches. Most of these hours are broadcast or read speech, and much of the audio is Modern Standard Arabic (MSA). Several of the biggest sets are licensed for non-commercial use only. Conversational Gulf speech, contact-centre telephony and code-switched speech are the thinnest slices, and they are usually what a production system hears.
Arabic speech recognition datasets at a glance
"Hours" is the figure the source states; some corpora count raw audio and others count transcribed segments, so treat small differences as noise. "Light" transcription means existing captions or transcripts were aligned to audio automatically. "Full" means humans transcribed the audio segment by segment.
| Dataset | Hours (as stated) | Varieties | Style / channel | Transcription | Licence (as stated) |
|---|---|---|---|---|---|
| QASR | ~2,000 | Mostly MSA; Gulf, Levantine, North African, Egyptian; code-switching | Al Jazeera broadcasts | Light | "Non-Commercial Purpose ONLY" (card) |
| MGB-2 | ~1,200 | More than 70% MSA, rest dialectal | Al Jazeera, 2005–2015: conversations, interviews, reports | Light | Challenge data from the organisers; terms not stated in the paper |
| MASC | 1,000 | Multi-regional, multi-dialect | YouTube, 700+ channels, 16 kHz | See card | CC BY 4.0 |
| SADA | ~667 | Najdi, Hijazi, Janubi, Shamali, Khaliji, MSA, plus Egyptian, Levantine, Iraqi, Yemeni, Maghrebi | 57+ Saudi TV shows; clean, noisy, music and car environments | Transcribed, with dialect and speaker labels | CC BY-NC-SA 4.0 |
| Common Voice 27.0 (ar) | 158 recorded, 93 validated | Mixed; no curated dialect labels | Read prompts, 1,679 volunteer voices | Prompt text | CC0-1.0 |
| FLEURS (ar_eg) | ~10 train per language | Egyptian locale | Read FLoRes sentences | Prompt text | CC BY 4.0 |
| Casablanca | ~48 | Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, Yemeni | TV series episodes on YouTube (URLs, not audio) | Full, plus gender, dialect, code-switching | CC BY-NC-ND 4.0; validation and test only (card) |
| MGB-3 | 16 | Egyptian | YouTube, seven genres | Four independent transcripts on dev data | Challenge data; terms not stated |
| MGB-5 | ~13 transcribed; ~3,000 dialect-ID | Moroccan (ASR); 17 countries (dialect ID) | YouTube | Full (ASR part); country labels only (dialect ID) | Challenge data; terms not stated |
| Mixat | 15 | Emirati mixed with English | Two public podcasts, conversational | Transcripts plus transliteration and translation | CC BY-NC-SA 4.0 (card) |
| Gulf Arabic CTS | ~2,800 minutes; 975 speakers | Gulf | Telephone | See catalogue | LDC User Agreement for Non-Members (fee) |
| CALLHOME Egyptian | 120 calls, up to 30 min each | Cairene Egyptian | Telephone, unscripted | See catalogue | LDC User Agreement for Non-Members (fee) |
| Arabic Speech Corpus | Small, single speaker | MSA, one speaker | Studio recordings, built for TTS | Phoneme-level | CC BY 4.0 |
Broadcast and MSA corpora
The big Arabic ASR datasets come from television news. MGB-2 released about 1,200 hours from 19 Al Jazeera programmes aired between 2005 and 2015, and its authors estimate that more than 70% of the speech is MSA. QASR extends the same archive to about 2,000 hours, with multi-dialect and code-switched segments and linguistically motivated segmentation. Both are lightly supervised: existing human transcripts, which can omit or correct what speakers actually said, were aligned to the audio automatically. The Casablanca paper flags this as a source of mismatch between transcript and speech.
SADA is the Gulf counterweight. It comes from Saudi television rather than a pan-Arab news channel, and its DataPort listing labels each segment's dialect, speaker type and recording environment. That makes it the most useful public set for Saudi varieties, as long as your use is non-commercial. MASC adds 1,000 hours crawled from more than 700 YouTube channels under a CC BY 4.0 card. Its card calls it multi-regional, multi-genre and multi-dialect; check its per-clip metadata before relying on it for any one variety.
Broadcast audio suits pre-training, media monitoring and captioning. It is a poor proxy for a customer calling a bank from a car: presenters speak clearly, one at a time, into studio microphones, and lean toward MSA.
Read and crowd-sourced speech
Read-speech sets give clean pairs of audio and known text. The Common Voice Scripted Speech 27.0 metadata lists 158 recorded and 93 validated hours of Arabic from 1,679 voices (collection cut-off 11 September 2026), released as CC0-1.0. FLEURS reads translated FLoRes sentences, with around 10 hours of training data per language. Its own card warns that read-speech results can differ from performance in noisier, real-world settings. The Arabic Speech Corpus is a single-speaker studio set built for speech synthesis, not recognition.
These sets suit smoke tests and licence-clean fine-tuning. None gives you curated dialect labels, and people do not talk to a voice assistant the way they read a sentence off a screen.
Dialectal and conversational corpora
This is where public data gets thin. Casablanca describes itself as the largest fully supervised Arabic dialect dataset: about 48 hours of human transcription across eight countries, with gender, dialect and English/French code-switching labels. It ships YouTube URLs rather than audio, because of copyright, so clips can disappear. Its authors also note that more than 60% of speakers are male in every dialect except Moroccan. MGB-3 (16 hours of Egyptian YouTube) and MGB-5 (about 13 transcribed hours of Moroccan) are smaller still.
Conversational code-switching has one notable public set: Mixat, 15 hours of Emirati speech mixed with English from two podcasts. Its authors built it because current speech recognition resources fall short on Emirati speech. Telephone speech lives mostly in the LDC catalogue: Gulf Arabic Conversational Telephone Speech (about 2,800 minutes from 975 speakers) and CALLHOME Egyptian (120 calls). CALLHOME was published in 1997 and the Gulf corpus was recorded in 2004; both are casual calls, not service interactions.
Dialect labels are another weak point. The Arab Voices project (January 2026), which harmonises 31 datasets across 14 dialects, reports that corpora define dialect by country, by coarse region or by ISO-style codes, which makes pooling them unreliable. It also notes that metadata such as speaker demographics is often limited or absent. Automatic dialect labels will not rescue you: in LabelBench, a local language model reached 22.4% exact-match accuracy across five dialect regions from text, on the fixed held-out manifest (chance is 20%). For more on building dialect labels that hold up, see Arabic dialect identification.
Here is why the style gap matters. The same request, "I want to check my balance", looks like this across the registers a banking voicebot will hear:
| Register | Utterance | Transliteration |
|---|---|---|
| MSA (broadcast, read) | أريد أن أتحقق من رصيدي | urīdu an ataḥaqqaq min raṣīdī |
| Najdi / Gulf | أبي أشوف رصيدي | abī ashūf raṣīdī |
| Hijazi | أبغى أشوف رصيدي | abghā ashūf raṣīdī |
| Egyptian | عايز أشوف رصيدي | ʿāyiz ashūf raṣīdī |
| Gulf with English | أبي أشيك الـ balance حقي | abī ashayyik il-balance ḥaggī |
A model trained mostly on the first row has seen the noun and little else. The verb of wanting, the verb of looking and the English insert change from row to row, and each also has several spellings (أبغى / ابغى / أبغا). Our guide to Arabizi and code-switching covers how to label the last row consistently.
Licences: what you can use commercially
Read the licence on the card, then the provenance. The licence says what the curators permit; provenance says whether they had the rights to permit it.
| Licence on the card | Datasets here | What it generally allows |
|---|---|---|
| CC0-1.0 | Common Voice | Public-domain dedication; no attribution required. |
| CC BY 4.0 | FLEURS, MASC, Arabic Speech Corpus | Commercial use with attribution. |
| CC BY-NC-SA 4.0 | SADA, Mixat | Non-commercial only; share-alike on adaptations. |
| CC BY-NC-ND 4.0 | Casablanca | Non-commercial, and no distribution of modified versions. |
| "Non-Commercial Purpose ONLY" | QASR | Research and other non-commercial use. |
| LDC user agreement | Gulf Arabic CTS, CALLHOME Egyptian | Terms set by the agreement and fee; read it for your use case. |
| Not stated | MGB-2, MGB-3, MGB-5 | Ask the challenge organisers before assuming anything. |
Three traps catch teams most often:
- Web-sourced audio. A CC BY card on a corpus crawled from YouTube or TV covers the curators' work. It does not transfer the broadcaster's or the uploader's rights in the audio. Casablanca's decision to ship URLs instead of files shows its authors took this seriously.
- Models trained on non-commercial data. Whether weights trained on NC data can be used commercially is a legal question the licence text does not settle. Get advice before a model trained on SADA or QASR ships in a product.
- Voices are personal data. Speech can identify a speaker, and your own call recordings fall under privacy law. See our PDPL data residency checklist before customer audio leaves your systems for transcription.
What is missing from public Arabic speech data
Map your deployment against this list. Every item you need is a gap public data will not close.
- Contact-centre telephony. None of the corpora above covers 8 kHz banking, telecom or government service calls, with hold music, cross-talk and account numbers read aloud.
- Conversational Saudi and Gulf speech on your channel. SADA labels Najdi, Hijazi and Khaliji, but it is television. The LDC Gulf telephone corpus was recorded in 2004. For the varieties themselves, see Saudi Arabic dialects for AI.
- Code-switching beyond one pairing. Mixat covers Emirati with English. Gulf English in Riyadh offices, Maghrebi French and Levantine English inserts have almost nothing public at conversational scale.
- Smaller varieties. Casablanca's authors single out Emirati, Yemeni and Mauritanian as under-represented, yet Casablanca releases only validation and test splits, and MGB-5's transcribed Moroccan set is about 13 hours.
- Balanced demographics. Gender, age band and region are recorded unevenly across corpora, and some sets skew heavily male.
- Domain vocabulary. Clinical dictation, legal proceedings and financial terms read in dialect are absent from public sets.
- Consistent spelling. Dialects have no official orthography. In MGB-3, four transcribers still disagreed on about 13% of words after normalising common letter variants. Mixing corpora mixes their conventions, which adds noise to both training labels and WER.
The effect shows up in evaluation. On SADA, the Open Universal Arabic ASR Leaderboard (December 2024 paper) found that every model tested did best on MSA and declined significantly on Egyptian and Khaliji. The authors attribute this to how heavily public training data leans toward MSA. Our piece on why dialect coverage sets your model's ceiling covers the same effect in text models.
When to commission custom Arabic speech data
Public corpora are the right start for pre-training, baselines and comparability. Commission custom collection or annotation when two or more of these are true:
- Your deployment dialect is not a labelled variety in any corpus whose licence fits your use.
- Your channel (phone, in-car, far-field, messaging voice notes) does not match studio or broadcast audio.
- You need commercial rights and the matching corpus is NC, ND or has unstated terms.
- Your users code-switch, and you need a written rule for how English or French inserts are transcribed.
- You need domain terms, names or numbers that public sets never contain.
- You need a held-out test set nobody has trained on, sliced by dialect, gender and age.
Even then, custom data rarely replaces public data. The usual pattern is to pre-train or start from a model built on broadcast hours, then fine-tune and evaluate on a smaller, well-specified in-dialect set. Write the transcription guidelines first, so the new hours share one convention. Set speaker quotas per variety and channel, not per "Arabic"; without them, a collection drifts toward whoever is easiest to record.
That second step is the work Bayanat Labs does. Our Arabic speech data collection and transcription is produced by vetted native speakers who pass dialect screening before they touch client audio. It covers transcription, timestamping, diarization, and accent and emotion tagging, calibrated against gold standards with adjudication and rework loops. Where you need broader dialect data collection across 25+ varieties, the same process applies. Hosting is in-region by default. A fixed-scope pilot with gold-standard QA and a quality report, usually within two weeks, lets you measure the data against your own test set before you scale.
- The big hours are broadcast. QASR, MGB-2, MASC and SADA account for most public Arabic speech, and news corpora are majority MSA.
- Licences split the field. Common Voice, FLEURS, MASC and the Arabic Speech Corpus are permissive on their cards. QASR, SADA, Casablanca and Mixat are non-commercial. Check the licence before commercial use.
- Conversational and telephone data is the gap. Public telephone corpora are small, old and casual; none of the corpora compared here covers contact-centre calls.
- Dialect labels don't pool cleanly. Corpora define dialect differently, so harmonise labels and spelling before you mix sets.
Frequently asked questions
What is the largest Arabic speech recognition dataset?
Among transcribed corpora, QASR is the largest at about 2,000 hours of Al Jazeera broadcasts, followed by MGB-2 at about 1,200 hours. Both use light supervision, so transcripts were aligned automatically rather than written fresh by humans. The MGB-5 dialect-identification release has about 3,000 hours, but it carries country labels, not transcripts. Size says little about fit: both are broadcast news, and both are majority MSA.
Is there a free Arabic speech dataset I can use commercially?
A few corpora carry permissive licences on their cards: Common Voice Arabic is CC0-1.0, and FLEURS, MASC and Halabi's Arabic Speech Corpus are CC BY 4.0, which allows commercial use with attribution. QASR, SADA, Casablanca and Mixat are non-commercial. Also check provenance: a licence on a YouTube-derived dataset cannot grant rights to audio the curators never owned. This is not legal advice; check the licence before commercial use.
Which Arabic speech dataset covers Saudi dialects?
SADA, released by SDAIA's National Center for AI with the Saudi Broadcasting Authority, is the main one: about 667 hours from more than 57 TV shows, labelled with Najdi, Hijazi, Janubi, Shamali and Khaliji among other varieties. It is licensed CC BY-NC-SA 4.0, and it is television audio, not phone calls. Our guide to Saudi Arabic dialects for AI covers the varieties themselves.
Are there Arabic call-centre or telephone speech datasets?
Only a handful, and they are old. The Linguistic Data Consortium distributes conversational telephone corpora such as Gulf Arabic Conversational Telephone Speech and CALLHOME Egyptian Arabic, under LDC user agreements with a fee. They are casual calls, not customer-service calls. None of the corpora compared here contains banking, telecom or government contact-centre audio, so teams building for those channels usually collect or annotate their own.
How should I benchmark an Arabic ASR model on dialects?
Start with public test sets for comparability. The Open Universal Arabic ASR Leaderboard uses SADA, Common Voice, MASC and MGB-2. Then build a small held-out set from your own channel, sliced by dialect, gender and age band. Apply one orthographic normalisation to both references and hypotheses (alef, yaa and taa marbuta variants) before computing WER. Otherwise you are measuring spelling conventions, not recognition.