500 hours of Arabic smartphone speech
Arabic speech recorded on smartphones in 2025–2026. Scripted and spontaneous subsets have separate labels so your team can choose, combine, and measure them deliberately.
What this dataset is for
For teams training mobile speech recognition or assistants that need to understand Arabic recorded on a phone.
Commercial license; permitted training use and delivery scope agreed in the license.
What’s included
- 500 hours of Arabic smartphone speech
- Collected in 2025–2026
- Separate labels for scripted and spontaneous subsets
- Alignment, evaluation, and agent-trajectory sets available to license with the audio
Full specifications & collection details
- Count
- 500 hours
- Language / dialect
- Arabic; ask us for the register and dialect breakdown.
- Speakers
- Count and composition available on request
- Collection year
- 2025–2026
- Training task
- Mobile speech recognition and spoken-assistant training
- Source
- Smartphone recordings; scripted and spontaneous subsets labeled separately.
- License
- Commercial license; permitted training use and delivery scope agreed in the license.
What to check in your sample
- Speakers, dialects, and hours in each subset
- Devices, recording conditions, transcripts, and alignment
- Companion dataset versions and any links to audio records
We’ll send the available sample and documentation so your team can check the fit before licensing.
Common questions
Can we distinguish scripted speech from spontaneous speech?
Yes. The subsets are labeled separately. Request the hour split and sample records to choose the balance that fits your training or evaluation.
What other data can we license with the audio?
Arabic alignment, evaluation, and agent-trajectory sets are available alongside the audio. Ask for their versions, formats, and any record-level links; they are not assumed to be included in the audio license.