Skip to content
Bayanat Voice · Available to license

500 hours of Arabic smartphone speech

Arabic speech recorded on smartphones in 2025–2026. Scripted and spontaneous subsets have separate labels so your team can choose, combine, and measure them deliberately.

What this dataset is for

For teams training mobile speech recognition or assistants that need to understand Arabic recorded on a phone.

Commercial license; permitted training use and delivery scope agreed in the license.

What’s included

  • 500 hours of Arabic smartphone speech
  • Collected in 2025–2026
  • Separate labels for scripted and spontaneous subsets
  • Alignment, evaluation, and agent-trajectory sets available to license with the audio
Full specifications & collection details
Count
500 hours
Language / dialect
Arabic; ask us for the register and dialect breakdown.
Speakers
Count and composition available on request
Collection year
2025–2026
Training task
Mobile speech recognition and spoken-assistant training
Source
Smartphone recordings; scripted and spontaneous subsets labeled separately.
License
Commercial license; permitted training use and delivery scope agreed in the license.
What to check in your sample
  • Speakers, dialects, and hours in each subset
  • Devices, recording conditions, transcripts, and alignment
  • Companion dataset versions and any links to audio records

We’ll send the available sample and documentation so your team can check the fit before licensing.

Common questions

Can we distinguish scripted speech from spontaneous speech?

Yes. The subsets are labeled separately. Request the hour split and sample records to choose the balance that fits your training or evaluation.

What other data can we license with the audio?

Arabic alignment, evaluation, and agent-trajectory sets are available alongside the audio. Ask for their versions, formats, and any record-level links; they are not assumed to be included in the audio license.