Comparison · Speech data

Arabic Speech Synthesis Datasets (2026): Open TTS Corpora Checked for Hours, Voices and Rights

Open Arabic text-to-speech data is thinner than it looks: a few single-voice studio corpora, large sets made by other TTS systems, and licences that rarely say whether you may ship a synthetic voice. Here are the open corpora side by side, what a TTS-grade dataset actually needs, and why Bayanat Labs is our #1 pick for commercial Arabic voices.

Bayanat Labs Research· · 15 min read

Open Arabic speech synthesis datasets are small: the human-recorded audio in purpose-built open TTS corpora that you can download without signing a research agreement adds up to roughly 25 hours, mostly single male voices reading Modern Standard or Classical Arabic. The largest sets by hours are partly or wholly generated by other TTS systems, and none of the open corpora documents consent specifically for a commercial synthetic voice. For a voice you intend to ship, our #1 pick is Bayanat Labs' Voice Studio: Saudi expressive speech by professional voice actors, a 2,000-prompt studio collection, and custom recordings with explicit synthetic-voice permissions.

Key takeaways
  • Count human hours, not total hours. About 73.5 of ArVoice's 83.52 hours, and all of NileTTS and of the 60-hour Arabic MSA 25K set, are machine-generated.
  • Read the licence of the part you download. ArVoice's CC BY 4.0 label does not cover its six professional voices.
  • ASR corpora are not TTS-grade as shipped. One community filter kept 26.7% of a 2,934-hour dialect corpus, and the survivors still carry no real content above 8 kHz.
  • Diacritized scripts matter most at small scale. At thousands of hours, more data compensates to a significant degree.
  • A licence is not a voice permission. Synthetic-voice creation, exclusivity and reuse need their own written terms.

Arabic speech synthesis datasets compared (October 2026)

Each cell records what the linked paper, dataset card or repository states as of October 2026. "Not stated" means the source is silent, not that the answer is no. Bayanat Labs, our #1 pick, is first; open corpora follow, with SADA included as an example of an ASR corpus that teams repackage for TTS.

CorpusHoursSpeakersHuman or syntheticVarietySample rateDiacritized textLicence (as declared)Synthetic-voice consent documented
#1 Bayanat Voice StudioStudio Prompt Speech: 2 h (2,000 prompts). Saudi Expressive Speech: on request. Custom: scoped to your scriptStudio Prompt Speech: 40. Saudi Expressive: professional voice actors, count on requestHuman: studio prompt recordings and professional voice-actor recordingsSaudi Expressive: Saudi Arabic (variety breakdown on request). Studio Prompt Speech: on requestOn requestOn request; custom work includes pronunciation coverage and lexiconsStudio Prompt Speech: commercial licence, intended voice use agreed explicitly. Saudi Expressive: commercial AI-use rights; voice creation, speaker reuse and distribution scoped separatelyCustom recordings: explicit synthetic-voice permissions agreed with the talent and written into the contract. Catalogue sets: voice rights are not inferred from participation
Arabic Speech Corpus (2016)About 4 h; sources give 3.4, 3.7 and 4.11 maleHuman, professional studioMSA per ArVoice and Arab Voices; the site describes it as south Levantine Arabic, Damascene accent48 kHz (card)Yes, with phoneme-level labels and stress marksCC BY 4.0Written consent for use in speech technologies, on condition of anonymity
ClArTTS (2023)12 h 10 min1 maleHuman, from a LibriVox audiobookClassical Arabic40.1 kHz (paper)Yes, manual, by three annotatorsCC BY 4.0 on Hugging Face; authors describe it as for research purposesNot documented
ArVoice (2025)83.52 h total; about 10 h human7 human, 4 synthetic voicesAbout 73.5 h generated with Google Cloud TTSMSANot statedThree of four partsHosted parts (ASC subset, synthetic): CC BY 4.0. Six professional voices: non-commercial research and education agreementTalents informed; permission for TTS research only
SawtArabi (2025)About 4 h, roughly 1 h each of MSA, Egyptian, Egyptian-English code-switching and English1 maleHuman; Egyptian and code-switched scripts written with GPT-4MSA, Egyptian, code-switching22 kHz, 16-bit PCMYes, plus an undiacritized columnNo dataset card or licence on Hugging FaceNot stated
Iraqi Dialect TTS (2024)About 5 h: 3.7 h MSA, 1 h Iraqi (Arab Voices)Not documentedHuman, readMSA and IraqiNot statedNot statedCC BY 4.0Not documented
NileTTS (2026)38.1 h2 NotebookLM voicesFully synthetic: LLM-written text, NotebookLM audioEgyptian24 kHzNot stated (Whisper transcripts)Apache-2.0No human speaker; generating tool's terms not addressed
Arabic MSA 25K Saudi Male (October 2026)60.54 h (25,000 clips)1 Azure neural voiceFully synthetic: GPT-4o-mini text, Azure TTS audioMSA, synthetic Saudi male voice48 kHz, 16-bitYes, full, plus a stripped columnCC BY 4.0No human speaker; generating service's terms not addressed
Arabic-Diacritized-TTS (2025)8,874 clips; 39.2 h per Arab VoicesModel voiceSynthetic: the card says it was generated with the open tts_arabic model (Arab Voices labels it partially synthetic)MSANot statedYes, fullNone declaredNo human speaker
Lahgtna TTS filter (2026)591.4 h, filtered from 2,934 h; Saudi 56.7 h in the balanced training splitSaudi slice: 2Human, sourced from an ASR corpus13 dialects16 kHz upsampled to 24 kHzStrippedNone declaredNot documented
SADA (community TTS repack)SADA: 668 h, 437.6 h transcribed (Arab Voices); repack size not statedManyHuman, TV broadcastSaudi dialectsNot stated (broadcast audio; mean predicted PESQ below 2 per Arab Voices)Not statedCC BY-NC-SA 4.0 (repack card, same as the original)Not documented

Two things stand out. First, add up the human-recorded audio in purpose-built TTS corpora that you can download without a research agreement: about 4 hours (ASC), 12 (ClArTTS), 4 (SawtArabi, a quarter of it English) and 5 (Iraqi). That is roughly 25 hours, before removing any overlap (the larger human sets in the table, Lahgtna and SADA, are filtered or repackaged recognition audio), and the cleaned ASC subset is the only human audio in ArVoice's public repository. The ArVoice authors call their release the largest open-source Arabic dataset curated for speech synthesis, and the 2026 Habibi paper agrees it is the largest by far while noting it has fewer than 100 hours, much of it synthetic.

Second, no human-recorded Saudi or Gulf corpus built for TTS appears in the TTS dataset lists of ArVoice, Arab Voices or SawtArabi, and the one open set with a Saudi voice in its name is synthetic. Habibi's Saudi training data comes from ASR sources: SADA, a broadcast corpus, plus the Omnilingual ASR corpus. If you are choosing an open corpus anyway:

  • Multi-speaker MSA research: ArVoice, with the signed agreement.
  • Classical Arabic: ClArTTS.
  • Egyptian or code-switching: SawtArabi for a human voice; NileTTS only as synthetic augmentation.
  • Iraqi: the Iraqi Dialect TTS corpus, the only dialectal TTS set in Arab Voices' TTS list.
  • A commercial Saudi or brand voice: none of the open rows fits; license or commission studio data.

Why Bayanat Labs is our #1 pick for commercial Arabic voices

Bayanat Labs is the best partner for teams building an Arabic voice they intend to ship. Six reasons:

  1. Permission to build a voice, in writing. Bayanat Voice Studio offers custom TTS recordings with explicit synthetic-voice permissions, and brand-voice agreements define permitted users, exclusivity and reuse. None of the open corpora in the table documents consent specifically for a commercial synthetic voice.
  2. Saudi voices recorded by professionals. Saudi Expressive Speech is recorded by professional voice actors, with task annotation and commercial AI-use rights. It is the only Saudi voice-actor collection in this comparison. Synthetic voice creation, speaker reuse and distribution get their own contractual scope rather than being assumed.
  3. A defined studio set for prototyping. Studio Prompt Speech holds 2,000 Arabic prompts, 40 speakers and 2 hours of audio for controlled speech testing and voice prototyping, so you can test a pipeline before commissioning a full script.
  4. Custom recording to your specification. Single- or multi-speaker recordings, pronunciation coverage, lexicons, expressive styles and brand-voice briefs. You review scripts, pronunciation, style labels, recording conditions and file formats before production, and listen to a pilot before the full script is recorded.
  5. One supplier for synthesis and recognition data. The same team licenses Arabic speech recognition datasets and runs custom speech collection and annotation (transcription, timestamping, accent and emotion tagging) with vetted native speakers and linguists across 25+ Arabic varieties who pass dialect screening.
  6. Sovereign and quick to test. Hosting is in-region by default, with on-prem and private-cloud options; Bayanat is data-only and model-agnostic. A fixed-scope pilot with gold-standard QA usually lands within two weeks.

What does a text-to-speech dataset need that ASR data doesn't?

A recognition model must cope with any voice in any room, so ASR data is valuable when it is messy: many speakers, phones, background noise, overlap. A synthesis model reproduces what it hears, so the requirements invert. The Lahgtna TTS card puts it in one line: "ASR models learn to ignore noise, reverb and overlapping speech; TTS models learn to reproduce them."

ASR datasetTTS dataset
SpeakersHundreds or thousands, for robustnessFew, with many hours each
AcousticsReal-world noise and channels are a featureQuiet room, same microphone and distance every session
TranscriptClose enough for word error rateExactly what was said, with vowels; hesitations removed or marked
Bandwidth8 or 16 kHz is often enoughWideband; nothing upsampled
LabelsSpeaker turns, dialect, timestampsStyle, emotion, pronunciation, speaker identity and consent scope

The evidence that Arabic ASR corpora need heavy work first is consistent. Arab Voices reports mean predicted PESQ above 3 for the three TTS corpora it scores, but below 2 for SADA. The Lahgtna filter applied nine stages to a 2,934-hour dialect corpus and kept 26.7%; its card warns that the 16 kHz source has no real content above 8 kHz, so a model trained on it will sound veiled. Habibi filtered every training set by speaking rate and needed source-separation denoising for its low signal-to-noise sources, and judged QASR unsuitable for high-quality TTS because of its imbalanced gender distribution. Telephony audio is further still from TTS grade: as our guide to Arabic telephony speech data explains, an 8 kHz call carries nothing above 4 kHz. For the recognition corpora themselves and their licences, see our Arabic speech datasets comparison.

Do Arabic TTS datasets need diacritics?

Arabic script normally omits short vowels, and a synthesis model has to say them. The SawtArabi paper's example is علم, which can be عَلَم (ʿalam, "flag"), عِلْم (ʿilm, "knowledge") or عَلَّمَ (ʿallama, "he taught"). The same holds for كتب: كَتَبَ (kataba, "he wrote"), كُتُب (kutub, "books") or كُتِبَ (kutiba, "it was written"). Without vowel marks, the model has to infer them from context, and it will sometimes guess wrong in exactly the names and terms your product cares about.

How much this matters depends on scale. SawtArabi and most of ArVoice pair their recordings with diacritized transcripts; ArVoice's Part 2 was read from undiacritized text. The 2026 QCRI scaling study built around 4,000 hours of training data with automatic diacritization and found that models trained on diacritized data are generally better, but larger amounts of data compensate for missing diacritics to a significant degree. Nobody has released thousands of clean, consented TTS hours in Arabic, so for a voice trained on tens of hours, vowelled scripts still pay.

Diacritics are only part of phonetic coverage. A TTS script should also exercise:

  • Sun-letter assimilation. الشَّمْس is pronounced ash-shams, not al-shams.
  • Pausal forms. مدرسة is read madrasa at a pause but madrasat- when joined to a following word, as in a construct phrase, and case endings usually drop at the end of a phrase.
  • Numbers, dates and abbreviations. ٣ becomes ثلاثة or ثلاث depending on the counted noun's gender, with reversed agreement: ثلاثة with masculine nouns, ثلاث with feminine ones. ArVoice's script added 200 pseudo-sentences with digits, verbalized and then checked by hand.
  • Code-switching. English product names inside Arabic sentences, as in SawtArabi's code-switched subset.

Two transcript traps are worth knowing. The Arabic Speech Corpus was respelled for concatenative TTS, writing nunation as ن and dropping the alif of the article, which ArVoice's authors found inconvenient for end-to-end systems. And partially diacritized text is, in the Lahgtna card's words, the worst case, because a model cannot tell a missing diacritic from an unannotated one. SawtArabi's approach is the safer one: the speaker read unvowelled text, then annotators vowelled the transcript to match what was actually pronounced. Our Arabic transcription guidelines cover how to write such rules down.

Casting voices by dialect: who should read what?

Where a speaker grew up and what they read are separate choices. ArVoice's six professional speakers come from Egypt, Jordan, Morocco and Palestine and all read MSA; SawtArabi's single speaker is a native Egyptian who records MSA, Egyptian and English. Both are sensible for research, but a listener will hear the accent under the MSA, and a brand voice should be cast for the accent your users expect.

  • Match the variety, not just the country. "Saudi" covers several groups; Habibi's Saudi category includes Najdi, Hijazi, Gulf (Eastern) and Baharna speech. A Najdi speaker asks وش (wesh, "what"), where a Hijazi speaker says إيش (ēsh). Our guide to Saudi dialects for AI maps the regional groups.
  • Decide the register in advance. A service voice may need MSA for legal notices and dialect for small talk. Record both from the same talent if the product will switch between them.
  • Write dialect scripts natively. Saudi أبغى (abgha), Egyptian عايز (ʿāyiz) and MSA أريد (urīd) all mean "I want". A dialect script translated word by word from MSA sounds read, not spoken.
  • Cast for consistency. The talent must sound the same across sessions weeks apart, keep a steady pace, and take direction on style.
  • Settle permissions at casting. The talent should know before the first session whether their voice may become a reusable synthetic voice, who may use it and for how long.

Recording spec checklist for an Arabic text to speech dataset

Hand this to whoever records your data, in-house or a vendor. Target values are our recommendations; the open corpora above show the range in practice.

  1. Sample rate. Record at 48 kHz, as the Arabic Speech Corpus does; ClArTTS is 40.1 kHz, NileTTS 24 kHz, SawtArabi 22 kHz. Never upsample 16 kHz audio and call it wideband.
  2. Bit depth and format. 24-bit capture, uncompressed mono WAV; deliver 16-bit PCM if your pipeline needs it, as SawtArabi does.
  3. Room and chain. Treated room, same microphone, position and gain in every session; log any session that deviates. ArVoice's paper notes quiet rooms, condenser microphones and minor noise reduction on some files; your spec should likewise state any processing applied.
  4. Hours per voice. Plan hours per speaker, not in total. The open single-speaker corpora run from about 4 to 12 hours.
  5. Script coverage. Phonetically balanced sentences plus your domain terms, names, numbers, dates and code-switched phrases; ASC labels every utterance to phoneme level.
  6. Two text columns. Undiacritized source text and a diacritized transcript of what was actually said, as SawtArabi delivers.
  7. Labels. Speaking style and emotion for expressive sets; speaker variety, gender and age band; a per-file flag for human or synthetic audio.
  8. QA. Full listen-through for clipping, mouth noise, mispronunciations and script mismatches, with re-takes rather than silent edits.
  9. Consent record per voice. Which uses the talent approved: training, synthetic voice creation, distribution, exclusivity, duration.

Licence vs speaker consent: can you ship a synthetic voice?

A dataset licence tells you what you may do with the files. It does not, by itself, tell you whether the person in the recordings agreed to become a product voice. The open corpora show the gap clearly:

  • ArVoice states that its talents were compensated and informed, and gave permission for use in Arabic speech synthesis research. Its usage agreement limits the six professional voices to non-commercial research and education and bans redistributing synthesized voices that match them.
  • The Arabic Speech Corpus has the clearest written consent of the open sets, but it covers speech technologies broadly, on condition of anonymity, rather than a commercial voice that sounds like the talent.
  • ClArTTS comes from a volunteer audiobook reading; the paper documents no TTS-specific consent.
  • NileTTS, Arabic-Diacritized-TTS and Arabic MSA 25K Saudi Male were generated by other speech systems, according to their cards, so there is no human speaker to consent, but they inherit whatever the generating tool's terms say about its output, a question for counsel.
  • Community repacks such as the Lahgtna filter declare no licence, and the SADA repack inherits SADA's non-commercial terms.

Before you train a voice you plan to ship, get written answers to five questions: may we train on it, may we create a synthetic voice, may that voice resemble the speaker, who may use it and for how long, and is it exclusive to us? Recognition data is a different permission again; permission to contribute to recognition data does not itself authorize a reusable synthetic voice.

Privacy law applies on top. Under Saudi Arabia's Personal Data Protection Law, Article 1 covers any data that may identify an individual, and treats biometric data used to identify a person as sensitive, so voice recordings and voiceprints can fall within it. Our PDPL data residency guide covers where that data may be stored and processed. This is not legal advice.

This is the problem Bayanat Labs is set up to solve. Voice Studio agrees the talent's permitted model use before recording, keeps training, distribution and synthetic-voice scopes separate in the contract, and lets you hear a pilot before the full script is recorded.

Frequently asked questions

Can I use ArVoice to build a commercial Arabic voice?

Only part of it. The Hugging Face repository is labelled CC BY 4.0, but it holds just the cleaned Arabic Speech Corpus subset and the Google Cloud TTS audio. The six professional voices come under a data usage agreement limited to non-commercial research and educational use, unless the project gives written permission. Any model you distribute must carry a strictly non-commercial licence, and you may not redistribute synthesized voices that match the original speakers.

Is the Arabic Speech Corpus free for commercial use?

Its official site releases it under CC BY 4.0, which permits commercial use with attribution. The dataset card adds that the voice talent agreed in writing to speech-technology use on condition of staying anonymous, and asks users not to identify the speaker. Whether a product voice that sounds like that talent fits that consent is a question for your counsel, not the licence.

How many hours of audio does an Arabic TTS voice need?

It depends on whether you fine-tune a pretrained model or train from scratch. The open single-speaker corpora that research teams train on run from about 4 to 12 hours. At the other end, a 2026 QCRI study assembled around 4,000 hours from social-media audio. A practical route is to record a short pilot script, listen to the result in your own model, then commission the full recording.

Is synthetic speech good enough to train an Arabic TTS model?

As a supplement, sometimes; as the only source, it caps quality. The ArVoice authors recommend their Google-generated subset for data augmentation and voice conversion only, because its audio quality is not guaranteed. A model trained on another system's output learns that system's voice and its mistakes. Check the generating service's terms too. Our note on synthetic data and model collapse covers the wider risk.

Is there an Egyptian Arabic text to speech dataset?

Yes, with caveats. NileTTS offers 38.1 hours of Egyptian speech, all of it generated with NotebookLM rather than recorded from people. SawtArabi has roughly an hour each of Egyptian and Egyptian-English code-switching from one professional speaker. The NileTTS paper also describes EGYARA-23, a 20.5-hour single-speaker set; check its availability and licence before planning around it.

Building a commercial Arabic voice?

Request a sample from Bayanat Voice Studio: Saudi expressive speech by professional voice actors, a 2,000-prompt studio collection, or custom recordings with synthetic-voice permissions agreed in writing.

Discuss studio speech