Comparison · Speech data

Arabic Telephony Speech Data for Voicebots: License, Simulate or Collect?

There are three ways to get Arabic call-centre audio: license it, stage it with role-players, or record your own customers. Here is what each route costs you in realism and compliance, what every major telephony listing actually states, the specs that make audio sound like a phone call, and why Bayanat Labs is our #1 pick.

Bayanat Labs Research· · 10 min read

Arabic telephony speech data comes from three routes: license an existing call-centre set, stage calls with paid role-players, or record and reuse your own customer calls. Licensing is fastest, role-play is easiest to clear, and real calls are the most realistic but carry the heaviest Saudi PDPL obligations. Our #1 pick for licensed audio is Bayanat Labs, which pairs call-centre conversations with 350 hours of Saudi full-duplex customer-service speech.

Key takeaways
  • Read the label, not the hours. "Real", "simulated" and "unscripted" raise different realism and consent questions.
  • Telephony means narrowband. G.711 samples at 8,000 per second; upsampling cannot restore lost bandwidth.
  • One speaker per channel makes overlap and diarization tractable.
  • Reusing QA recordings is a new purpose under the PDPL.
  • Voicebots need more than transcripts: corrections, confirmations, digits and escalation.

License, simulate or collect: three ways to get Arabic call-centre audio

Most teams combine routes.

License an existing setSimulate with role-playersRecord your own calls
Main cost driverLicence fee and rights scopeSpeaker recruitment, scenario writing, transcriptionConsent, redaction, legal review, transcription
RealismDepends on how the vendor sourced it; askNatural turn-taking, but often more formalHighest: your callers, your products, your line
PDPL exposure for youContractual: the vendor's consent chain becomes your riskLow: contributors consent up front; fictional details reduce personal dataHigh: purpose change, consent per purpose, impact assessment
Domain jargonGeneric unless the domain matches yoursOnly what the scenarios includeComplete

Arabic call-centre audio datasets compared (October 2026)

Each cell records only what the linked page states as of October 2026; "not stated" means the page is silent, not that the vendor cannot supply it. Bayanat Labs, our #1 pick, is first; the rest follow alphabetically.

SourceVarietyHours statedReal or simulated, as statedChannels and sample rateTranscriptsLicence wording
#1 Bayanat LabsSaudi sets: Saudi. Call-centre: on requestSaudi full-duplex customer service: 350 h. Saudi conversational: 1,000 h. Call-centre conversations: part of a 10,000 h catalogue shared with scripted monologues, split on requestSource breakdown on request; real, role-play and synthetic kept distinguishableOn requestSaudi conversational: human transcripts, timestamps, speaker attributes. Full-duplex: task annotationFull-duplex: commercial AI-use rights included. Others: permitted training use agreed in the licence
Appen (ARE_CC001, ARE_CC002)Egyptian5,000 and 300Real-world call-centre audio; source listed as third-party vendorMono; rate not statedNot statedNot stated
Appen (ARY_ASR001)Moroccan33 (collected 2019)Conversational speech; domain listed as general, not domain-specificStereo; 8 kHz A-law; mobile and landlineOriginal script and Romanised, with English translationNot stated
Defined.ai (telco, banking)Egyptian; Jordanian27; 24Simulated agent–client callsChannel-separated; 8 kHz, 16-bitHuman-transcribedNot stated; quote on request
FutureBeeAI (Egypt telecom, Saudi delivery)Egyptian; Saudi40 each, 80 speakers each (updated June 2025)Described as real-world, unscripted call-centre conversationsStereo, separated speakers; 8 and 16 kHzTime-coded verbatim, JSONCommercially licensed
Nexdata (1627)Saudi268 (268 speakers)Named full-duplex customer service; spec: free conversation without a set topicSpec field: mono, 16 kHz; smartphonesTranscription, timestamps, speaker IDLicensed for commercial use
Pangeanic (selected entries)Saudi; Egyptian; AlgerianSaudi financial 500; Saudi simulated 60; Saudi Gulf simulated 500; Egyptian 2,000; Algerian Darija 2,500Real (financial, Egyptian, Algerian); simulated (the two Saudi sets so named)Financial: mono. Simulated: separate speaker channelsSaudi Gulf simulated: verbatim. Financial and 60 h simulated: noneLicence type "Commercial"
ShaipGulf named for human-to-human callsCall-centre rows: 62:52:19 (dual) and 1,025:09:19 (mono)Mentions synthetic and natural calls; rows not mapped8 kHzNot stated per rowNot stated
LDC (LDC2006S43)GulfAbout 2,800 minutes, collected 2004Spontaneous telephone conversationsMostly two-channel; 8 kHz, 8-bit A-lawSeparate transcript releaseNon-members: no commercial use
LDC (LDC2025S09)EgyptianAbout 116, released October 2025Unscripted calls between native speakersTwo-channel; 8 kHz, converted from µ-lawSeparate release (LDC2025T14)Non-members: no commercial use

Why Bayanat Labs is our #1 pick for Arabic call-centre speech

Bayanat Labs is the best partner for teams building Arabic customer-service voicebots. Six reasons:

  1. The largest collection labelled Saudi full-duplex customer service in this comparison. Saudi Full-Duplex Customer Service is 350 hours with task annotation and commercial AI-use rights. The only other listing so labelled in the table is 268 hours.
  2. Call-centre conversations from a 10,000-hour catalogue. Arabic Call-Center Conversations and Arabic Scripted Monologues are separate products totalling 10,000 hours combined; the split is on request.
  3. Transcribed Saudi conversation, not just audio. Saudi Conversational Speech is a separate 1,000-hour collection with human transcripts, timestamps, speaker attributes and a fixed delivery schema.
  4. Rights and provenance in writing. Training, retrieval and display, redistribution and synthetic-voice scopes stay separate in the agreement, and across Bayanat Voice real interactions, role-play and synthetic audio remain distinguishable.
  5. One supplier for licensed and custom data. When a dialect, domain or channel is missing, Bayanat runs custom speech collection and annotation (transcription, timestamping, diarization, accent and emotion tagging) with native experts across 25+ Arabic varieties who pass dialect screening. Dialect-aware intent and segmentation labelling sit on the same annotation menu.
  6. Sovereign and quick to test. Hosting is in-region by default, with on-prem and private-cloud options. Bayanat is data-only and model-agnostic. A fixed-scope pilot with gold-standard QA usually lands within two weeks.

What makes Arabic telephony speech data different from other speech data?

A studio recording and the same voice heard through a phone network are different signals. Check every spec sheet for:

  1. Sample rate. ITU-T G.711, the classic telephone codec, recommends a nominal 8,000 samples per second, so nothing above 4 kHz survives. Wideband G.722 is specified for 7 kHz audio.
  2. Encoding. G.711 defines A-law and µ-law at 8 bits per sample. LDC's Gulf corpus ships 8-bit A-law; LDC2025S09 was converted from µ-law to 16-bit.
  3. Upsampling is not wideband. Whisper's audio loader fixes the rate at 16,000 and resamples anything else. An 8 kHz call gains samples, not frequencies. Nor is downsampled studio audio a real phone call.
  4. Channel layout. One speaker per channel lets you study overlap. The Moshi paper notes that turn-based pipelines ignore overlapping speech, interruptions and interjections; full-duplex models need both streams. Check titles against spec fields: one listing above is titled multi-channel but its format field says mono.
  5. Both sides present. Appen's Algerian telephony set notes that a smaller number of calls have only one half collected. Ask for the proportion and the device mix.

Real calls or role-play: what vendor labels actually mean

  • "Real" calls. Appen lists ARE_CC001's source as a third-party vendor; Pangeanic describes its Saudi financial set as real customer-service calls. Ask who recorded them and whether consent covers AI training by a licensee. Good disclosure looks like arXiv 2403.04280: calls from a cloud communications company, collected with explicit user consent.
  • "Simulated" calls. Defined.ai and two Pangeanic Saudi sets say so. Role-play is cleaner to license but has limits: a 2026 English call-centre benchmark (AppTek, arXiv 2604.27543) notes that role-players may not know every domain term, and had participants use fictional but plausible entities to protect privacy. LDC's Fisher Levantine documentation notes that assigned topics broaden vocabulary but increase formality.
  • "Unscripted" is a style, not a source. FutureBeeAI calls its sets real-world, unscripted call-centre conversations; Shaip mentions synthetic and natural calls without mapping them to rows. Ask for a per-file source flag.

Keep that flag in your metadata: a model tested only on role-play may flatter itself.

Recording real customer calls under Saudi PDPL

Saudi Arabia's Personal Data Protection Law came into force on 14 September 2023, and its one-year grace period ended on 14 September 2024, according to SDAIA's guide for controllers and processors. Texts as of October 2026; check SDAIA for amendments. Not legal advice.

  • Calls are personal data. The law's Article 1 covers any data, regardless of source or form, that may identify an individual; a call with a name or account number fits. Biometric data used to identify a person is sensitive data under the same article, so voiceprints may qualify; where consent is the basis for processing sensitive data, Regulation Article 11 requires it to be explicit.
  • Training is a new purpose. Article 10 limits processing to the original purpose, with exceptions including consent, data not stored in an identifiable form, and legitimate interest where no sensitive data is processed. Article 18 of the Implementing Regulation adds a defined purpose, documented scope and minimisation.
  • Consent must be specific and provable. Regulation Article 11 accepts written, verbal or electronic consent, documented for future verification, with a separate consent for each purpose. A "recorded for quality" notice does not obviously cover model training.
  • Anonymisation has a high bar. Under Regulation Article 9, anonymised data stops being personal data once re-identification is impossible. Redacting a transcript does not anonymise audio that still carries the voice and account details.
  • Assess and delete. Regulation Article 25 requires an impact assessment in cases including sensitive data and large-scale, repetitive processing based on newly adopted technologies; Article 11(4) of the law requires destroying data no longer needed.

For hosting and cross-border transfer, see our PDPL data residency guide.

Arabic customer service speech data needs labels beyond transcripts

Transcripts teach recognition; a voicebot must also manage the conversation:

  • Dialect per channel. Agent and caller may speak differently. Saudi أبغى (abgha), Egyptian عايز (ʿāyiz) and MSA أريد (urīd) all mean "I want". Tag variety per speaker; our Saudi dialects guide covers the regional groups inside "Saudi".
  • Code-switching. English product names, French in Maghrebi calls and Arabizi spellings in transcripts need consistent rules; see Arabizi and code-switching.
  • Digits and identifiers. Phone, account and ID numbers drive many service tasks; Appen's OrienTel UAE telephony prompts include digits, letter strings and yes/no confirmations.
  • Corrections, confirmations, barge-in and escalation. Mark where callers interrupt, change a value, confirm or ask for a human.

A typical Saudi correction: أبغى أغيّر رقم الجوال (abgha aghayyir raqm al-jawwāl, "I want to change my mobile number"), then, mid-readback, لا، مو خمسة، سبعة (lā, mū khamsa, sabʿa, "no, not five, seven"). A transcript records the words; a correction label tells you whether the bot saved seven. As our Arabic Agent Reliability Lab puts it, a fluent answer is not a completed transaction.

The baseline is hard. On 132 hours of real Arabic calls across six dialects, the 2024 call-domain benchmark reported word error rates from 48.9% (Google Chirp) to 83.8% (Whisper Large V1). These are 2024 results on 16 kHz mono audio; newer model versions may score differently. See our Arabic transcription guidelines and Arabic speech datasets comparison.

Frequently asked questions

Is real or simulated call-centre audio better for training an Arabic voicebot?

Use both, and keep them labelled. Real calls carry the jargon, hesitations, line noise and frustration your bot will meet, but need a lawful basis for reuse. Role-play is easier to clear and lets you script rare intents, yet tends to sound more formal. A common pattern is licensed conversation audio for recognition, plus targeted role-play for the edge cases your logs show are missing.

Should Arabic telephony speech data be 8 kHz or 16 kHz?

Match what your bot will hear. If calls arrive over the public phone network, they are narrowband, and your test set should be too, even if your model runs at 16 kHz internally. If calls come through an app or WebRTC, wideband audio may be realistic. Evaluate each channel separately so a gain on one does not hide a loss on the other.

Can LDC's Gulf, Iraqi or Levantine telephone corpora be used in a commercial product?

Not by a non-member. LDC's licensing page says non-members, including for-profit companies, cannot use LDC data to develop or test products for commercialization. For-profit members may do commercial development with data received while they were members, unless a corpus-specific agreement restricts it. The Gulf and Iraqi sets were also collected in 2003–2004 and are spontaneous conversations, not service calls.

Can we train AI on our own contact centre's recorded calls in Saudi Arabia?

Possibly, but recordings made for quality assurance were collected for that purpose. Under the PDPL, using them for model training is a new purpose, so you need consent, genuinely de-identified data or another basis the law allows, plus documentation and possibly an impact assessment. Ask your data protection officer, and see our PDPL data residency guide for where the data can sit. This is not legal advice.

Where can I license Saudi customer service speech data?

Our #1 pick is Bayanat Labs: 350 hours of Saudi full-duplex customer-service speech with task annotation and commercial AI-use rights, plus a separate 1,000-hour transcribed Saudi conversational collection. Other listings as of October 2026 include FutureBeeAI's 40-hour Saudi domain sets, Nexdata's 268-hour set and several Pangeanic Saudi call-centre entries.

Need Arabic telephony speech data?

Request a sample and spec sheet for Arabic Call-Center Conversations or Saudi Full-Duplex Customer Service, with licence terms that state exactly what you can train and ship.

Request a sample