200,000 MSA sentences for named entity recognition
Annotated sentences for finding entities in Modern Standard Arabic. Review the entity labels and span conventions to see how they fit your extraction system.
What this dataset is for
For teams building Arabic entity extraction into search, document processing, or language-understanding systems.
Commercial license; permitted training use and delivery scope agreed in the license.
What’s included
- 200,000 Modern Standard Arabic sentences
- Annotations for named entity recognition
- A commercial license for the agreed use
Full specifications & collection details
- Count
- 200,000 sentences
- Language / dialect
- Modern Standard Arabic
- Collection year
- Available on request
- Training task
- Named entity recognition and entity extraction
- Source
- Annotated MSA text; source details available on request.
- License
- Commercial license; permitted training use and delivery scope agreed in the license.
What to check in your sample
- Entity types, span rules, and tokenization
- Text sources and annotation review process
- Export format and training/evaluation split boundaries
We’ll send the available sample and documentation so your team can check the fit before licensing.
Common questions
Is this dataset Modern Standard Arabic or dialectal Arabic?
It is Modern Standard Arabic. If your system needs regional dialects or mixed Arabic-English text, discuss that coverage as a separate custom annotation request.
Which entity labels and annotation format are used?
The label taxonomy and export format are available on request. Check the sample for entity types, nested or overlapping spans, and tokenization before mapping it to your model.