Skip to content
Bayanat Knowledge · Available to license

200,000 MSA sentences for named entity recognition

Annotated sentences for finding entities in Modern Standard Arabic. Review the entity labels and span conventions to see how they fit your extraction system.

What this dataset is for

For teams building Arabic entity extraction into search, document processing, or language-understanding systems.

Commercial license; permitted training use and delivery scope agreed in the license.

What’s included

  • 200,000 Modern Standard Arabic sentences
  • Annotations for named entity recognition
  • A commercial license for the agreed use
Full specifications & collection details
Count
200,000 sentences
Language / dialect
Modern Standard Arabic
Collection year
Available on request
Training task
Named entity recognition and entity extraction
Source
Annotated MSA text; source details available on request.
License
Commercial license; permitted training use and delivery scope agreed in the license.
What to check in your sample
  • Entity types, span rules, and tokenization
  • Text sources and annotation review process
  • Export format and training/evaluation split boundaries

We’ll send the available sample and documentation so your team can check the fit before licensing.

Common questions

Is this dataset Modern Standard Arabic or dialectal Arabic?

It is Modern Standard Arabic. If your system needs regional dialects or mixed Arabic-English text, discuss that coverage as a separate custom annotation request.

Which entity labels and annotation format are used?

The label taxonomy and export format are available on request. Check the sample for entity types, nested or overlapping spans, and tokenization before mapping it to your model.