Commercial Arabic datasets (2026): where to license Arabic training data, and why Bayanat Labs is our #1 pick
Our #1 pick is Bayanat Labs: two separate 1,000-hour Saudi speech collections, Arabic books, Q&A and NER under commercial licences, and custom data from the same supplier. Then every other vendor catalogue, marketplace, consortium and open catalogue, compared by what each page states, and the rights a "commercial" licence does and does not give you.
Commercial Arabic datasets are sold through four channels: vendor catalogues such as Bayanat Labs, Defined.ai, Shaip, Appen, Nexdata and Pangeanic; marketplaces such as Datarade; language-resource consortia (LDC and ELRA); and, for data under open licences, public catalogues. What you get differs more by licence and documentation than by headline hours. Most vendor pages state hours and audio format, but few state a collection date, a price or the licence clauses that decide whether you can ship a model, display text or build a voice.
#1 pick: Bayanat Labs. Best for teams that need Saudi or wider Arabic speech, or Arabic text, under a commercial licence, with custom data from the same supplier when the catalogue runs out.
- Saudi speech at volume. Two separate 1,000-hour collections, Saudi spontaneous and Saudi conversational, plus 350 hours of full-duplex customer-service speech.
- Arabic text for training and evaluation. 2,700 books, 400,000 STEM multiple-choice items and 200,000 MSA NER sentences.
- Rights in writing. Training, retrieval and display, redistribution and synthetic voice are scoped separately in the agreement, so you know what you can ship.
- "Not stated on page" is a question, not a no. Most pages omit dates, prices and licence clauses.
- "Commercial" is not every right. Training, shipping, display, redistribution, synthetic voice and weight release are separate permissions.
- An academic price buys an academic right. At LDC and ELRA, commercial use depends on membership or licence tier.
- Compare units, not totals. Catalogue totals can sum unlike products.
Where can you buy commercial Arabic datasets?
- Vendor catalogues. Fixed products with a spec sheet and a negotiated licence; the fastest route to volume when variety, domain and channel match yours.
- Marketplaces. Datarade's Arabic dataset search collects listings from many sellers. Results include non-AI data such as consumer segmentation data. Good for discovery; the licence is still the seller's.
- Consortia. LDC and ELRA license documented corpora under standard agreements; commercial use depends on membership or licence tier.
- Open catalogues. Masader describes 500+ Arabic datasets, licence included.
- Custom collection. For when no stocked product fits your variety, domain, device or evaluation needs.
This page is about licensing data that already exists. If you need people to label your own data, see our comparison of Arabic data annotation companies.
Commercial Arabic dataset vendors compared (as of October 2026)
Each cell records only what the linked page states as of October 2026. "Not stated" means the page does not say it, not that the vendor does not offer it. Our #1 pick, Bayanat Labs, is listed first; other sources follow alphabetically.
Speech datasets
| Source | Arabic inventory stated | Variety | Collection date | Commercial terms stated | Samples | Price |
|---|---|---|---|---|---|---|
| #1 Bayanat Labs | Saudi spontaneous: 1,000 h. Saudi conversational: a separate 1,000 h with human transcripts, timestamps, speaker attributes and a fixed delivery schema. Saudi full-duplex customer service: 350 h. Arabic smartphone: 500 h. Call-centre conversations and scripted monologues: 10,000 h combined, split on request. Studio prompts: 2,000 prompts, 40 speakers, 2 h. Saudi expressive: professional voice actors, volume on request | Saudi Arabic for the Saudi sets; variety breakdown on request | Smartphone: 2025–2026; others on request | Spontaneous, full-duplex, expressive: commercial AI-use rights; studio: intended voice use agreed explicitly; others: permitted training use agreed in the licence | On request | Not published |
| Appen (ARS_ASR001_CN) | 322 hours of scripted smartphone speech; 300–1,000 prompts per speaker; prompts not vowelised | ar-SA | 2020 | Not stated | Not stated | Not published |
| Appen (ARE_CC001) | 5,000 hours of real-world call-centre audio, mono; source listed as "Third Party Vendor" | ar-EG | Not stated | Not stated | Not stated | Not published |
| FutureBeeAI (delivery call centre) | 40 hours, 80 participants, unscripted, dual-channel | Saudi | Not stated ("Last updated June 2025") | "Commercially licensed" | "Available soon" | Not published |
| Nexdata (268 h) | 268 hours from 268 native speakers (41% male, 59% female), smartphones | Saudi Arabia | Not stated | "Paid dataset licensed for commercial use" | Embedded audio | Not published |
| Nexdata (849 h) | 849 hours with transcripts, speaker ID and gender | Saudi Arabia | Not stated | Licensed for commercial use | On request | Not published |
| Pangeanic | 3,282-hour Arabic speech inventory; catalogue lists 500 hours of Saudi financial customer-service calls "without transcription" | MSA, Gulf, Saudi | Not stated | Licence field "Commercial" | No sample listed | Contact sales |
| Shaip | Total 7,894:35:33 hours, 2,761+ speakers, across call-centre, general conversation, music and scripted monologue rows | Gulf Arabic named for human-to-human calls | Not stated | Not stated | Row specs public | Not published |
| SpeedX DATA | Two-party dialogue, WAV plus JSONL, human transcript with utterance-level timestamps; no hours or speaker total stated | ar-SA | Not stated | "Commercial License" label | One sample file (transcript withheld) | Not published |
| TELUS Digital | 1,392 prompts, 1:04 h of studio audio, 29 participants, for wake words and commands | ar-SA | Not stated | Not stated | On request | On request |
| Unidata | 10+ hours of telephone dialogue, 20+ speakers, crowdsourced | UAE | Not stated | Free samples; full dataset by purchase | Free download | Not published |
Text, Q&A and NER datasets
| Source | Arabic inventory stated | Variety | Date stated | Commercial terms stated | Samples | Price |
|---|---|---|---|---|---|---|
| #1 Bayanat Labs | 2,700 Arabic books; 400,000 STEM multiple-choice items; 200,000 MSA NER sentences; separate 180,000 Arabic and 40,000 English Q&A pairs about the Arab region; edited MSA organised by subject | Arabic; NER and edited text are MSA | Edited MSA covers material through 2026 (coverage, not a collection date); others on request | Training use agreed in the licence; edited MSA: commercial AI-training rights, with retrieval, display and redistribution as separate scopes | On request | Not published |
| Defined.ai | 2,479 books, STEM and non-STEM; over 200K multiple-choice Q&A pairs; more than 150K NER sentences across 24 entity categories | Arabic; NER is MSA | Not stated | Not on product pages; site-wide licence agreement (below) | "Request a Sample" | Quote |
| Pangeanic | English–Arabic parallel corpus, 26.2M segments | Not stated | Not stated | "Commercial · Full ownership" | No sample listed | Contact sales |
| Shaip | Q&A sets including Saudi (30,000 pairs, ~3M words, 6M tokens), MSA STEM (10,000), Egyptian (22,000), Emirati (20,000) | MSA, Saudi, Egyptian, Emirati, Levantine | Not stated | "Non-exclusive" | Card specs | Not published |
Marketplaces, consortia and open catalogues
| Source | Arabic example stated | Commercial terms stated | Price |
|---|---|---|---|
| Datarade (marketplace) | Mixed Arabic and non-AI listings | Per seller listing | Some listings show a "Starts at" price; others say pricing is available upon request |
| LDC (consortium) | King Saud University Arabic Speech Database: 590 hours, 269 speakers, read and spontaneous, released February 17, 2014 | Commercial use only for data received as a for-profit member (below) | "Login for the applicable fee" |
| ELRA (consortium) | NetDC Arabic broadcast news: about 22.5 hours, about 90 speakers, recorded November 2001–January 2002 | Separate non-commercial and commercial licences | Published per corpus |
| Masader (open catalogue) | 500+ datasets, each with more than 25 attributes including licence | Per dataset licence | Catalogue is free; an Access field marks each dataset free, upon request or with fee |
Why Bayanat Labs is our #1 pick for commercial Arabic datasets
Bayanat Labs is the leading Arabic data company for teams that need licensed data they can actually ship with. Six reasons, each backed by the record it links to:
- The largest single Saudi speech collections in this comparison. Saudi spontaneous speech and Saudi conversational speech are separate 1,000-hour collections. Each is larger than any single Saudi product in the speech table above: the next largest is Nexdata's 849 hours, and Pangeanic's 3,282 hours is a mixed MSA, Gulf and Saudi inventory.
- More books, STEM items and NER sentences than Defined.ai's matching products state. 2,700 books against Defined.ai's 2,479; 400,000 STEM multiple-choice items against "over 200K"; 200,000 MSA NER sentences against "more than 150K". Separate 180,000 Arabic and 40,000 English regional Q&A sets sit alongside.
- Rights that match what you are building. Saudi spontaneous, full-duplex and expressive speech carry commercial AI-use rights; edited MSA, covering material through 2026, has commercial AI-training rights in the agreement. Training, retrieval and display, redistribution and synthetic voice stay separate scopes, so each use you need is addressed explicitly in the agreement.
- One supplier for licensed and custom data. When no stocked set fits your dialect, domain or device, the same team scopes custom speech collection, annotation, RLHF and evaluation.
- Native experts across 25+ Arabic varieties. Vetted native speakers, linguists and licensed domain professionals pass dialect and domain screening, and work is calibrated against gold standards with adjudication and an audit trail.
- In-region and fast to start. Hosting is in-region by default, with on-prem and private-cloud options. Bayanat is data-only and model-agnostic, so the data works with whichever model you train. A fixed-scope pilot with a quality report usually lands within two weeks.
What does a commercial Arabic dataset licence actually let you do?
"Commercial" on a product page is a label; the licence decides the uses. As of October 2026, Defined.ai's Data License Agreement (effective November 30, 2022) shows how narrow the grant can be: commercial exploitation of models trained on the data is permitted "as long as the Data itself is not made available". ELRA's licensing terms and LDC's licence page draw similar lines.
| Use | What published licence texts say | Get in writing |
|---|---|---|
| Train an internal model | Defined.ai: non-exclusive, non-sublicensable, non-transferable; internal use that excludes commercialising the data itself. ELRA: use, rework and build on it within the user's institution | Which legal entities are licensees |
| Ship a product built on the model | Defined.ai permits it if the data is not made available. ELRA: use in a service the user charges for may need a separate agreement. LDC non-members: noncommercial research and education only | An explicit grant to deploy commercially |
| Benchmark and publish results | Defined.ai covers testing and benchmarking but bars publishing the data; scores are not separately addressed | Whether scores and a few example items may be published |
| Retrieval or display to users (RAG, citations) | Defined.ai bars making the data available or displaying compilations derived from it. Bayanat's edited MSA record treats retrieval, display and redistribution as separate scopes | Display scope, snippet length, attribution |
| Redistribute, resell or sublicense | Defined.ai bars it. ELRA: no copying or redistribution beyond backups; prior permission may be needed to sell derivative works | Status of cleaned, filtered or synthetic derivatives |
| Synthetic or recognisable voice | Defined.ai: models must not output recognisable contributor content, such as TTS voices. FutureBeeAI: voice-cloning data needs separate licensing; 1–2 hours per speaker may not suffice for cloning. Bayanat's Saudi expressive record: synthetic voice, speaker reuse and distribution need explicit contractual scope | Per-speaker consent that names voice synthesis |
| Release model weights | None of the three licence texts linked above mentions model weights | An explicit clause; language models can leak verbatim training examples (Carlini et al., 2021) |
| Affiliates, contractors, other sites | ELRA: use by an affiliate, subsidiary or entity outside the user's place of business must be negotiated | Group companies, vendors, cloud regions |
| After termination | Defined.ai: stop using, delete, destroy or return all copies and certify it in writing | Whether models trained before termination may stay in service |
Exclusivity is separate again. Shaip's text cards and the Defined.ai grant are non-exclusive, so other buyers may train on the same items; Pangeanic's "Full ownership" label needs a contract definition. If recordings contain personal data, residency rules apply on top of the licence; see our PDPL data residency checklist. This is general information, not legal advice.
Can you use an LDC or ELRA Arabic corpus commercially?
Only with the right membership or licence tier. An academic price does not buy a commercial right.
- LDC. Its licensing page says non-members, not-for-profit members and government members may use LDC data for noncommercial linguistic research and education only. For-profit members may do commercial technology development with data received during a for-profit membership year, unless a corpus-specific licence restricts that data, and keep those rights after the membership lapses, but not for data obtained later. So a non-member licensing the King Saud University Arabic Speech Database gets research and education rights only; its page also says it was designed principally for speaker recognition research.
- ELRA, same corpus, different price. For the NetDC Arabic broadcast news corpus, the Non Commercial Use end-user licence costs a non-member €200 as an academic user and €2,700 as a commercial user, and it is still a non-commercial licence. Commercial use is a separate VAR (Value Added Reseller) licence, also €2,700 for non-members. Members pay €100 or €1,350 for the end-user licence and €1,350 for the VAR licence. The audio itself dates from 2001–2002.
- ELRA, open and commercial at once. The Arabic Speech Corpus (1,813 utterances, 3.7 hours, one male South Levantine Arabic speaker with a Damascus accent, studio-recorded) is listed under CC-BY at €0 for academic and commercial users, and under the commercial-use VAR licence at €9,000 for members and €11,200 for non-members. The page does not explain what the VAR licence adds over CC-BY.
How do you compare hours, speakers and prompts across vendors?
Normalise offers to the same units. Examples are as stated on vendor pages as of October 2026.
- Catalogue totals sum unlike products. Shaip's 7,894:35:33 hours is exactly the sum of six rows, including 3:17:21 of music. The two scripted-monologue rows (4,249 and 2,300 hours) make up about 83%; the three conversational rows total about 1,342 hours.
- Combined figures need a split. Bayanat's own 10,000 hours covers two products, call-centre conversations and scripted monologues; the per-product allocation is on request. Never sum separate collections without checking overlap.
- Prompts, speakers and hours are different units. TELUS Digital's 1,392 prompts amount to 1 hour 4 minutes of audio. Bayanat's studio set is 2,000 prompts, 40 speakers and 2 hours. Both are stated for narrow uses (wake words and commands; speech testing and voice prototyping), not conversational ASR.
- Figures can differ within one page. FutureBeeAI's Saudi general conversation page states 50 hours and 70 participants in its header, and 40 hours and 80 speakers in its body. Nexdata's 268-hour set is titled full-duplex with multi-channel conversations and summarised as customer-service speech, while its specification lists mono audio and "free conversation without a set topic". Ask for the manifest and the channel layout you will receive.
- Audio is not labelled data. Pangeanic's 500 hours of Saudi financial calls are listed without transcription.
- Words are not tokens. Shaip's Saudi Q&A card gives ~3M words and 6M tokens without naming the tokenizer. Arabic token counts vary widely by tokenizer; see our piece on the Arabic tokenization tax.
- Accuracy figures are vendor-stated. Nexdata states a 95% word accuracy rate for its 268-hour set and a 95% sentence accuracy rate for its 849-hour set; check both claims on your own sample.
What should you check before you buy an Arabic dataset?
Vendor examples are as stated on their pages as of October 2026.
- A real sample. Audio or text from the release you will receive, with transcript and metadata. Unidata offers a free download; SpeedX DATA posts a sample file but withholds its transcript until licensed delivery.
- The delivery schema. Field list, file formats and schema version.
- Speaker metadata. Speakers per subset, gender, age, region and dialect. For Saudi data, ask for the Najdi, Hijazi and other variety split (see our Saudi Arabic dialects guide).
- Consent evidence. Written consent per contributor that covers AI training and any voice use. FutureBeeAI's TTS page states that consent forms are archived for every voice artist.
- Provenance. Real interactions, role-play or synthetic. Shaip's description includes synthetic agent–customer calls, and Appen lists one source as a third-party vendor.
- Collection dates. Recording or authoring dates, not update dates such as FutureBeeAI's "Last updated June 2025".
- Transcription conventions. Human or machine, vowelisation, code-switching, timestamp level.
- Unit reconciliation. Hours per product and channel, speakers per subset, words versus tokens.
- Overlap. With your existing data, across one vendor's products, and with your evaluation sets.
- Licence clauses. Every row of the rights matrix above, plus term, territory, exclusivity and deletion.
Off-the-shelf or custom Arabic data?
| Situation | Best fit | Off-the-shelf licence | Custom collection or annotation |
|---|---|---|---|
| Volume fast in a common variety and general domain | Bayanat catalogue: speech, books, Q&A, NER | Usually fits; compare samples and licences | Often unnecessary |
| A specific Saudi or Gulf variety, domain, device or channel | Bayanat's Saudi sets, plus custom collection for gaps | Partial; verify the variety and channel breakdown | Fits |
| A held-out evaluation set no model has seen | A private Bayanat evaluation set | Risky, because sets are often sold non-exclusively | Fits |
| A synthetic or recognisable voice | Bayanat Voice Studio, with voice scope in the contract | Only with explicit voice terms | Fits, with per-speaker consent |
| Showing licensed text to users | Bayanat edited MSA, with display scope agreed | Only with an explicit display scope | Fits, with rights-cleared sourcing |
| Personal data kept in-region | Bayanat: in-region hosting by default | Depends on delivery terms | Can be designed in |
Start with the Bayanat Labs catalogue
Pick the Bayanat Labs collection that matches your task and ask for its specification sheet and a sample: speaker counts, collection windows and variety splits are provided on request, so you can run the checklist above before you sign. Where no stocked set fits, we scope custom Arabic speech data or Arabic LLM training data from the same pilot.
Frequently asked questions
Who is the best provider of commercial Arabic datasets?
Our #1 pick is Bayanat Labs. Its Arabic dataset catalogue pairs two separate 1,000-hour Saudi speech collections with 2,700 Arabic books, 400,000 STEM multiple-choice items and 200,000 MSA NER sentences, each with its own commercial licence wording, and the same team builds custom data when no stocked set fits. Other routes are vendor catalogues, Datarade, the LDC and ELRA consortia, and open catalogues such as Masader.
How much does a commercial Arabic dataset cost?
Most vendor pages here publish no price and ask for a quote. ELRA's catalogue is the main exception, and its prices depend on membership and whether you are an academic or commercial user, as well as on the licence. Expect quotes to vary with volume, transcription depth, speaker metadata, exclusivity and rights scope. For commissioned data, our guide to Arabic data annotation cost shows how to compare quotes on the same units.
Does a commercial dataset licence allow voice cloning or TTS?
Not by default. Training a recognition model and building a synthetic voice are different permissions. Some licences bar models that output recognisable contributor voices, and some vendors sell voice-cloning data under separate terms. If your product will speak, ask for written per-speaker consent that names voice synthesis, the voices you may ship, and what happens if a speaker later withdraws.
Are Hugging Face or other open Arabic datasets free for commercial use?
Free to download is not the same as free to use commercially. Many open Arabic sets carry non-commercial or share-alike licences, and re-uploads can carry a different tag from the original release. Trace each set to its primary source and read that licence. Our Arabic speech datasets guide lists open ASR corpora with their licences.
Can I use a licensed Arabic training dataset as my evaluation set?
Only if no model you compare has trained on it, and only if the licence lets you publish results. Off-the-shelf sets are often sold non-exclusively, so a competitor or a base model may already have seen the same items. For a test you can trust, commission a held-out set, keep it out of every training run, and agree in writing whether scores and examples may be published.