Service · Speech & audio

Arabic speech data collection & transcription for AI

Custom recordings and transcripts across 25+ Arabic varieties: scripted or spontaneous, speaker-balanced, and transcribed to a written convention you can score against.

Arabic speech data for ASR and voice AI has to sound like your users, not like the news. Bayanat Labs collects recordings from vetted native speakers across 25+ Arabic varieties. We then transcribe, timestamp, diarize and tag them to a written convention agreed with you. The work is for teams building or evaluating speech recognition, voice assistants and call-center analytics. Data is hosted in-region by default, and every engagement starts with a fixed-scope pilot, usually within two weeks.

The public numbers make the case for custom data. On the Saudi SADA test set, the Open Universal Arabic ASR Leaderboard (December 2024) reports that the best-scoring model it tested, NVIDIA's Conformer-CTC large with a language model, reached 19.2% word error rate (WER) on MSA but 48.2% on Khaliji. Whisper-large-v3 went from 28.0% to 59.9%. Casablanca (EMNLP 2024) tested eight dialects and measured zero-shot Whisper-large-v3, before text pre-processing, at 48% WER on Jordanian up to 87% on Mauritanian. In-dialect audio is how you close that gap, and it rarely exists for your variety on your channel. For more on this, see why dialect coverage sets your model's ceiling.

25+Arabic varieties and dialects
In-regionhosting by default; on-prem and private cloud available
~2 weekstypical fixed-scope pilot

What we deliver

Speech collection

Scripted, semi-scripted and spontaneous recordings from screened native speakers, against speaker quotas and specs agreed up front.

Transcription

Arabic-script transcripts in verbatim and clean layers, tagging fillers, noise, overlap, code-switches and personal data.

Timestamps & segmentation

Utterance- or word-level timing, with segments sized to your training pipeline.

Speaker diarization

Who spoke when, with overlapping speech marked and agent and customer roles labelled on call audio.

Accent, dialect & emotion tagging

Per-speaker variety labels checked by native listeners, plus emotion, sentiment or intent at the segment level.

ASR evaluation sets

Held-out test sets sliced by variety, channel and speaker group, so WER is reported per slice and not as one blended average. The same principle applies when evaluating Arabic LLMs.

Planning Arabic speech data collection

"Arabic" is not a sampling frame. Name the varieties you deploy in, such as Najdi, Hijazi and Eastern Province rather than "Saudi". Then set quotas for gender, age band and channel within each one. Without quotas, a dataset drifts toward whoever is easiest to record. Casablanca's Palestinian subset, for example, is 92% male. On SADA, the leaderboard above also found that most models had higher WER for elderly speakers than for adults, and slightly higher WER for female speakers than for male.

  • Scripted (read prompts): wake words, commands, names, numbers, product terms and TTS. The result sounds read.
  • Semi-scripted (scenarios): prompts such as "dispute a card charge" or "book a clinic appointment". You get natural phrasing while still covering the domain you planned for.
  • Spontaneous (free conversation): the real mix of disfluency, overlap and code-switching. It gives the most robust ASR and costs the most to transcribe.

Record on your production channel (telephony, app, far-field or in-car) and fix sample rate, bit depth, channel layout and file format during scoping.

Transcription conventions for dialectal Arabic

Dialects have no official spelling, which makes your transcription convention part of your model. In the MGB-3 challenge, four transcribers spelled Egyptian Arabic however they saw fit. After normalizing common letter variants, they still disagreed on about 13% of words. That disagreement becomes noise in both your training labels and your WER. The CODA* guidelines, which cover 28 city dialects from Rabat to Muscat, are a sound base for a project lexicon. These are the decisions that matter most:

DecisionOption AOption BSensible default
Fillers, false startsVerbatim: يعني… اه… كنت، كنت بدّي روح عالبنكClean: كنت بدّي روح عالبنك ("I wanted to go to the bank", Levantine)Verbatim with tagged fillers for ASR; clean layer for analytics
Dialect soundsPhonetic: حالچ (ḥālich, Gulf feminine "how you are")Conventional: حالكConventional spelling plus lexicon; phonetic only for pronunciation work
Word boundariesمبيقولهاشما بيقولهاش ("he doesn't say it", Egyptian)Pick one, list it, normalize before scoring
English code-switchLatin, tagged: عندي meeting بكرةArabic script: عندي ميتنق بكرة ("I have a meeting tomorrow")Latin with a language tag; Arabic-script layer if output must be all-Arabic
NumbersWords: أبغى أحوّل ألفين ريالDigits: أبغى أحوّل 2000 ريال ("I want to transfer 2,000 riyals", Gulf)Words in verbatim (they carry inflection); digits in a normalized layer
Personal dataKept in a restricted layerTagged: رقم جوالي [PHONE]Tag in text, mute in shared audio

Code-switching into English or French is common in many dialects. On Mixat, a set of Emirati-English podcast recordings, existing models struggle with the mixed speech. Casablanca asked annotators for both a Latin-script and an Arabic-script rendering of each foreign word. Arabizi (chat Arabic in Latin script, such as 3ayez) belongs to text data. If you need it paired with speech, treat it as a separate output layer.

How an Arabic speech data project runs

  1. Scoping. We agree varieties, speaker quotas, channel, domain, consent wording, recording specs and delivery schema.
  2. Guidelines and gold set. We write the convention and dialect lexicon. Native linguists then transcribe and adjudicate gold clips, and contributors pass dialect screening and training against them.
  3. Pilot. A fixed slice is recorded and transcribed end to end, with a quality report broken down by variety.
  4. QA and adjudication. Second-pass listening, gold checks and adjudicated disagreements feed rework loops and guideline updates.
  5. Scale and delivery. Batches ship with audio, transcripts, timestamps, speaker metadata and an audit trail on every task. This is Vet → Train → Produce → QA → Deliver applied to audio.

Quality control for Arabic audio

We check audio for clipping, silence, the wrong channel and the wrong speaker before transcription starts. Each speaker's variety is confirmed by a native listener, not taken from self-report alone. Gold clips go into every transcriber's queue and are scored by WER after normalization to the convention. A share of clips is transcribed twice, independently, so agreement is measured rather than assumed. Disagreements go to adjudication and rework, and the ones that keep recurring become lexicon entries. Your quality report states the method and the results on your data. We do not quote generic accuracy figures.

When to automate Arabic transcription

Machine pre-transcription with human correction works when the model is already strong on your variety and channel, so test it on a gold slice first. At the dialect error rates above, correction turns into rewriting, and editors anchor on wrong drafts. Keep dialect metadata human-verified. In LabelBench, a local language model reached 22.4% exact-match accuracy on five dialect regions from text on the fixed held-out manifest (chance is 20%), and it called 75% of Gulf examples Egyptian. LabelBench's Safe Automation Coverage metric sets the rule: automate only the share of the queue that can be accepted while the 95% lower confidence bound stays above your accuracy target.

Consent, privacy and residency for voice data

Under Saudi Arabia's Personal Data Protection Law, a recording that can identify someone is personal data. Biometric data used for identification falls in the stricter sensitive tier. So consent wording and purpose are agreed before recording, and call audio gets personal-data tagging and redaction. Hosting is in-region by default, with on-prem and private-cloud options. Access is least-privilege and audit-logged, and all contributors work under NDA. Our PDPL and data residency checklist lists the questions to put to any speech data partner.

Use cases by industry

  • Telecom and government: IVR and voice-assistant ASR, tested per dialect.
  • Finance and banking: Arabic call-center transcription with diarization, PII tags and intent or sentiment labels.
  • Healthcare: clinical dictation and patient conversations, with terminology reviewed by licensed physicians.
  • Media and retail: captioning, search and voice commerce across dialects.

In-house, crowd platform or specialist vendor?

CriterionIn-house teamCrowd platformSpecialist vendor
Dialect verificationOnly varieties your staff speakUsually self-reportedScreened native speakers per variety
ConventionFull control, full upkeepGeneric; drifts across workersProject lexicon, maintained with you
QADepends on bandwidthSpot checks or majority voteGold sets, adjudication, rework
ResidencyYour systemsWorkers and storage anywhereContractual; in-region options
Best fitSmall, highly sensitive scopesLarge, simple read speechDialectal, conversational or regulated audio

Start with a pilot

Pick one or two varieties and one channel. The pilot has a fixed scope and covers collection or transcription, gold-standard QA and a quality report by variety, usually within two weeks. After that, you can scale as a managed, embedded or enterprise engagement. The same team handles text through Arabic data annotation and labeling and dialect corpora through Arabic dialect data collection. Tell us your markets and use case.

Frequently asked questions

How many hours of Arabic speech data do I need to fine-tune an ASR model?

There is no universal number. It depends on how far your base model is from your target variety and channel. Score the model on a small dialect-sliced test set, add data in batches and re-measure after each one. A few well-balanced hours in the right dialect often do more than many hours in the wrong one.

Can I use public Arabic speech datasets instead of custom collection?

For pre-training and benchmarking, often. On their own, rarely. Many public corpora are broadcast or read speech; MGB-2, for example, is Al Jazeera television audio, which is a long way from 8 kHz customer calls. Check each licence for commercial use, then fill the gaps you can measure with targeted collection.

Should Arabic transcripts include diacritics?

For most ASR training, no. Everyday Arabic is written without short vowels, so full diacritization adds cost and a new source of disagreement. TTS and pronunciation work are the exception. Some projects also keep selected marks such as tanween and hamza, as Casablanca did. Whatever you choose, write the rule into the convention.

What audio format and sample rate should Arabic speech data use?

Match the deployment channel. Telephony is narrowband, typically 8 kHz, so call-center models should train on real call audio, not downsampled studio recordings. App audio is usually 16 kHz or higher, and TTS needs studio capture. Deliver lossless WAV or FLAC, keep speakers on separate channels where possible, and log device and environment as metadata.

Scope an Arabic speech data pilot

Tell us your target varieties, channel and use case. We will propose a fixed-scope pilot with a written transcription convention, gold-standard QA and a quality report, usually within two weeks.

Scope a speech pilot