1,000 hours of Saudi spontaneous speech
Unscripted Saudi speech for models learning to recognize naturally phrased Arabic. Task annotation and commercial AI-use rights are included in the collection.
What this dataset is for
For ASR and spoken-language teams that need Saudi speech beyond read scripts.
Commercial AI-use rights included; permitted uses are specified in the agreement.
What’s included
- 1,000 hours of Saudi spontaneous speech
- Task annotation
- Commercial AI-use rights
Full specifications & collection details
- Count
- 1,000 hours
- Language / dialect
- Saudi Arabic; request variety breakdown
- Speakers
- Count and composition available on request
- Collection year
- Available on request
- Training task
- Spontaneous-speech recognition and spoken-language understanding
- Source
- Spontaneous speech; recording origin and speaker mix available on request.
- License
- Commercial AI-use rights included; permitted uses are specified in the agreement.
What to check in your sample
- Speaker count, Saudi language varieties, and collection dates
- Task labels, transcript coverage, and delivery format
- Published benchmark score, benchmark name, model, version, and metric
We’ll send the available sample and documentation so your team can check the fit before licensing.
Benchmark report
Ask for the benchmark name and version, model, metric, and score for this collection. The report can help your team judge whether the data fits your task.
Ask for the benchmark reportCommon questions
Is this the 1,000-hour Saudi Conversational Speech collection?
No. They are separate products, each containing 1,000 hours. This listing is Saudi Spontaneous Speech. The conversational listing has its own specifications, including human transcripts, timestamps, and speaker attributes.
Where can we review the benchmark result and speaker breakdown?
Ask for the product datasheet and published benchmark result. We can help you review the benchmark name, model, metric, and score alongside the speaker mix and collection dates.