Gulf Arabic dataset guide: Emirati, Kuwaiti, Qatari and Khaleeji data for AI
Public Khaleeji data is older, more Saudi and more MSA than its labels suggest. Here are the main text and speech sets with their dates and licences, the countries they miss, and how to fill the gaps without rework.
In short: the main public Gulf Arabic datasets are the Gumar text corpus (2016), LDC's Gulf telephone speech (collected 2004), the MADAR and QADI multi-dialect sets, SDAIA's SADA broadcast speech, and a handful of small recent speech sets such as Mixat (Emirati–English) and a 2025 Qatari corpus. Most are old, dominated by Saudi or MSA content, or licensed for research only. None of the sets listed here gives you recent, commercially usable, country-labelled Khaleeji speech for every Gulf state.
- "Gulf" is several dialects. Linguists define it narrowly; large datasets such as Gumar stretch it to all six GCC states.
- Saudi dominates the big sets. Saudi is the largest share of Gumar, and SADA is Saudi television.
- The data is dated. Apart from Saudi-focused SADA, large resources date from 2004–2021; recent Gulf-specific sets are 12–15 hours or a few hundred speakers.
- Kuwaiti, Bahraini and Omani speech are the clearest gaps. None of the public speech corpora listed here is specific to any of them.
- Code-switching is measured, not anecdotal. 36% of Mixat's utterances mix Emirati Arabic and English.
What counts as Gulf Arabic?
It depends on who is counting. In the linguistic literature summarised by Khalifa et al. (2016), Gulf Arabic strictly means the varieties of Kuwait, Bahrain, Qatar, the seven Emirates and Al-Hasa in eastern Saudi Arabia. The same paper notes that Omani, Hijazi, Najdi and Baharna Arabic are usually left out of Gulf Arabic grammars because they differ considerably. The Gumar project then widens the term to any indigenous variety of the six GCC states, and the multi-country sets below use that wider sense.
So a "Gulf" label can mean Kuwaiti coastal speech or Najdi from Riyadh, and a model trained on one is not automatically good at the other. For Najdi, Hijazi and other Saudi varieties, see our Saudi dialect dataset guide; this page covers the Gulf as a whole.
How do Emirati, Kuwaiti, Qatari and other Gulf varieties differ?
None of these features is uniform across a country.
| Feature | Example | Where it is reported |
|---|---|---|
| /q/ pronounced /g/ (or /dʒ/) | gāl "he said"; dʒidir "pot" | Certain Gulf dialects (Khalifa et al., 2016) |
| /k/ pronounced "ch" /tʃ/ | chaf "palm (of the hand)" | Some Gulf varieties (Khalifa et al., 2016) |
| /dʒ/ pronounced /j/ | ryāl vs rjāl "men" | Urban Qatari favours ryāl; Bedouin Qatari prefers rjāl (Bouamor et al., 2025) |
| Future with ba- or rāḥ | ba- prefix instead of MSA sa- | rāḥ as a future marker in Kuwaiti, Bahraini and some Saudi varieties (Khalifa et al., 2016) |
| Negation with mū | mū, mub, muhub, hub | Gulf Arabic generally, citing Holes (Khalifa et al., 2016) |
| Possessive māl / ḥag | li-ktāb māl… "the book of…" | Gulf Arabic generally (Khalifa et al., 2016) |
| Distinct feminine plural forms | gāmat "she stood up" vs gāman "they (f.) stood up" | Some Saudi, Emirati and Omani varieties (Khalifa et al., 2016) |
| Epenthetic -n- in active participles | māʿṭinhum "I've given them" | Some Emirati and Baharna varieties (Khalifa et al., 2016) |
First, spelling varies with pronunciation: the same word can be written with ق or ج, and ك or چ, depending on the writer. Second, Gulf varieties also share features with neighbouring dialects, which is one reason Arabic dialect identification struggles to separate Gulf from its neighbours.
Gulf Arabic datasets: text and speech
The table lists public resources that trace back to a primary source, with sizes and licences as each source states them, as of September 2026. On ACL Anthology and arXiv the licence shown covers the paper, not the data. This is not legal advice.
| Dataset | Varieties | Modality & size | Year | Licence | Link |
|---|---|---|---|---|---|
| Gumar | GCC-wide; documents labelled SA, AE, KW, OM, QA, BH | Text: 112.4M words from about 1,200 forum novels | 2016 | Not stated; web interface | LREC 2016 |
| Gulf Arabic diacritization set | Emirati (Dubai) subset of Gumar | Text: 19,850 diacritized words | 2022 | Open-source release planned; licence not named | WANLP 2022 |
| MADAR | Riyadh, Jeddah, Doha, Muscat among 25 cities | Text: 2,000 translated BTEC travel sentences per city; 12,000 for five cities including Doha | 2018 | Internal research and evaluation only | CAMeL Lab |
| QADI | 18 countries including all six GCC states | Text: 540,590 tweets | 2021 | Not stated in the README | GitHub |
| ZAEBUC | Learner essays by UAE university students; not dialect-labelled | Text: 602 Arabic and English essays by 397 students (about 33K Arabic and 88K English words) | 2022 | Described as publicly available | LREC 2022 |
| Gulf Arabic Conversational Telephone Speech (LDC2006S43) | "Gulf Arabic"; countries not specified | Speech: about 2,800 minutes, 976 call sides, 8 kHz | Collected 2004; released 2006 | LDC user agreement for non-members | LDC |
| SADA | Various Saudi dialects | Speech: about 667 hours of TV audio | 2023 (Kaggle update) | CC BY-NC-SA 4.0 | Kaggle |
| QASR | Mostly MSA; GLF listed among its dialects | Speech: about 2,000 hours of Al Jazeera, 2004–2015 | 2021 | Stated as available to the research community; licence not named | arXiv |
| Mixat | Emirati–English | Speech: 15 hours, 5,316 utterances from two podcasts | 2024 | Announced for research use | arXiv |
| ZAEBUC-Spoken | MSA, Gulf, Egyptian, English | Speech: 12 hours of Zoom meetings | 2024 | Described as publicly available; licence not named | arXiv |
| Casablanca | Eight dialects; Emirati is the only Gulf one | Speech: about 48 hours across eight dialects, with transcription, dialect and code-switching labels | 2024 | Research only; YouTube links and annotations, not audio | arXiv |
| Qatari intra-dialectal corpus | Qatari, Urban and Bedouin | Speech and transcripts: 255 speakers | 2025 | Described as publicly available | ArabicNLP 2025 |
Two cautions. LDC's licensing policy allows commercial technology development only for for-profit members, and only with data received while they were for-profit members, so licensing the Gulf telephone corpus as a non-member does not let you train a product on it. And Hugging Face re-uploads, such as a cleaned 2018–2020 Gulf tweet set whose card points to its source through a shortened link, cannot grant more than the upstream release allows. For a wider view of speech options, see our comparison of Arabic speech recognition datasets.
Which Gulf countries does the public data actually cover?
| Country | Text with a country label | Speech with a country label |
|---|---|---|
| Saudi Arabia | Gumar (largest share), MADAR (Riyadh, Jeddah), QADI | SADA |
| UAE (Emirati) | Gumar, QADI, Gulf diacritization set (Dubai) | Mixat, Casablanca |
| Qatar | Gumar, MADAR (Doha), QADI | Qatari intra-dialectal corpus |
| Kuwait | Gumar, QADI | None among those listed |
| Oman | Gumar, MADAR (Muscat), QADI | None among those listed |
| Bahrain | Gumar (smallest share), QADI | None among those listed |
Gumar's authors report that 92% of the corpus is Gulf Arabic, with Saudi the most dominant sub-dialect and Bahraini the least. The biggest speech set is Saudi too. In the speech sets listed here, Kuwaiti, Omani and Bahraini speakers could appear only under a generic "Gulf" tag, as in the LDC telephone corpus. This is why dialect coverage sets a model's ceiling.
The freshness gap in Gulf Arabic data
Most large Gulf resources describe the region's language as it was a decade or more ago. The LDC telephone calls were collected in 2004. QASR's Al Jazeera audio runs from 2004 to 2015. Gumar was released in 2016, and MADAR's sentences (2018) are translations of travel-domain BTEC sentences, not natural text. The newest large text set among those listed, QADI, was released in 2021.
Resources released since 2024, including podcasts and online meetings, are closer to current speech but small or narrow: 15 hours of podcasts in Mixat, 12 hours of role-play meetings in ZAEBUC-Spoken, one Gulf dialect in Casablanca, and 255 Qatari speakers. None is a general-purpose, commercially licensed training set. The one large recent speech set, SADA, is mostly Saudi TV audio under a non-commercial licence.
Broadcast data does not close the gap. In the MGB-2 test set described in the QASR paper, the variety split is 78% MSA and 22% dialectal, and no Gulf country is among the top five by dialectal segments. A good score there mostly proves a model handles news Arabic.
Code-switching in Gulf business Arabic
Recent corpora measure how often English enters Gulf Arabic speech. In Mixat, 1,947 of 5,316 utterances (36%) are code-switched; the authors describe bilingual Emiratis who often mix and switch between their dialect and English. The Qatari corpus reports systematic code-switching with English, most frequent among younger Urban women, while Bedouin speakers, especially older men, rarely code-switch.
An illustrative sentence of the kind a support agent in Dubai or Kuwait might hear (constructed, not taken from a corpus):
أبي أسوي booking بس الـ app ما يشتغل (abī asawwī booking bas il-app mā yishtaghil, "I want to make a booking but the app isn't working").
The hard tokens are the English nouns, the Arabic article الـ attached to an English word, and the Gulf verb أبي ("I want"). Decide before labelling whether English stays in Latin script, how article-plus-English tokens are tagged, and whether language tags are word- or span-level. Our guide to Arabizi and code-switching annotation covers each rule.
Annotation pitfalls specific to Gulf Arabic
- Accepting "Gulf" as a label. Record the speaker's country and, where possible, region and urban or Bedouin background. The Qatari corpus shows lexical and phonological differences inside one small country.
- Normalising away pronunciation. Writers choose چ or ك, ق or ج. Keep the raw form and add a normalised layer, or you lose the dialect signal you paid for. The Gulf diacritization guidelines are one reference point for a written standard.
- Trusting an LLM dialect filter. In LabelBench v0.1, a local language model reached 22.4% exact-match accuracy across five dialect regions (chance 20%) on the fixed held-out manifest. It labelled 75% of Gulf items Egyptian and got 4% of them right. The Arabic-focused Fanar 1 9B did better, at 47.8%, but a small supervised character model still beat it. Validate any automatic filter on a native-labelled gold set.
- Transcribing without a code-switching policy. Fix script, tagging and hesitation conventions before the first file, as set out in our Arabic transcription guidelines.
How to build a Gulf Arabic dataset: a decision checklist
- Name the countries and channels. A Kuwaiti banking voicebot and an Emirati retail chat need different speakers and audio.
- Use public data for evaluation, not as your training base. Gumar, QADI and MADAR suit baselines and dialect-ID tests; their research-only or unstated licences and their dates make them weak product foundations.
- Collect where the coverage table has no speech set. Kuwaiti, Omani and Bahraini speech need new recordings, with consent wording that covers training.
- Stratify within each country. Balance urban and Bedouin speakers, age and gender, as the Qatari corpus was designed to, and log them as metadata.
- Collect natural code-switching. Do not scrub English from prompts; measure its rate in your pilot.
- Keep provenance per item: source, licence, consent and collection date.
This is the work Bayanat Labs does. Gulf, including Saudi/Najdi, Emirati and Kuwaiti, is one of the dialect families we cover within 25+ Arabic varieties, alongside Levantine, Egyptian, Iraqi and Maghrebi. Contributors are vetted native speakers who pass dialect screening before they touch client data, and their work is calibrated against gold standards with adjudication and an audit trail on every task. Our Arabic dialect data collection and annotation service and our Arabic speech data and transcription service both start with a fixed-scope pilot with gold-standard QA and a quality report, usually within two weeks.
Frequently asked questions
What is the Gumar corpus?
Gumar is a Gulf Arabic text corpus from NYU Abu Dhabi, described by Khalifa et al. (LREC 2016). It holds about 112 million words from roughly 1,200 online forum novels, with a sub-dialect label per document (Saudi, Emirati, Kuwaiti, Omani, Qatari, Bahraini, or other). The text is browsable through a web interface. The paper does not state a dataset licence, so confirm terms with the authors before any training use.
Is there a Kuwaiti Arabic dataset?
Not as a large, dedicated public corpus with a primary source. Kuwaiti text appears as one class inside multi-dialect resources: the KW documents in Gumar and the Kuwait class in the QADI tweet set. None of the public speech corpora listed in this guide is specific to Kuwaiti. If Kuwait is a launch market, plan to collect Kuwaiti speakers and record their country and region as metadata.
Does MGB-2 or QASR contain Gulf Arabic?
Some, but both are mostly Modern Standard Arabic news from Al Jazeera. The QASR paper lists GLF among its dialects, yet the MGB-2 test set it describes is 78% MSA, and the top five countries by dialectal segments are Egypt, Syria, Palestine, Algeria and Sudan. Broadcast news is a poor proxy for Gulf call-centre or voice-assistant speech.
Is there a Gulf Arabic dataset on Hugging Face?
Yes, several user uploads, such as a cleaned 2018–2020 Gulf tweet set. That card carries a CC BY 4.0 tag but points to its Twitter source only through shortened links, and other uploads repackage existing corpora. A re-upload's tag cannot override upstream terms, so trace each one to its original release first.
Can I use SADA to train a commercial Gulf speech model?
Not under its public licence. SADA, released by SDAIA's National Center for AI with the Saudi Broadcasting Authority, is published on Kaggle under CC BY-NC-SA 4.0, which rules out commercial use. It is also mostly Saudi TV audio, so it says little about Emirati, Kuwaiti or Omani speech.