2,700 Arabic books for language-model training
Book-length Arabic text gives models examples of sustained arguments, longer narratives, and subject-specific language. Browse the titles and subjects before selecting the material you need.
What this dataset is for
For model teams that need longer Arabic texts for language modeling or deeper coverage of a particular subject.
Commercial license; permitted training use and delivery scope agreed in the license.
What’s included
- 2,700 books in the catalog
- Arabic long-form text
- A license covering the selected titles and uses
Full specifications & collection details
- Count
- 2,700 books
- Language / dialect
- Arabic; ask us for the register and dialect breakdown.
- Collection year
- Available on request
- Training task
- Language modeling and domain adaptation
- Source
- Arabic books; title and publisher list available on request.
- License
- Commercial license; permitted training use and delivery scope agreed in the license.
What to check in your sample
- Titles, publishers, editions, and subject coverage
- Document structure and available file formats
- Rights per title, duplicate handling, and volume before and after cleaning
We’ll send the available sample and documentation so your team can check the fit before licensing.
Common questions
Can I see which books and subjects are included?
Request the title and publisher list, subject breakdown, and a sample. Those details help you assess whether the collection fits your model’s intended subject coverage.
Does the book license also let us display or redistribute the text?
Permitted uses are set out in the license for your selected titles. Tell us whether you need training, retrieval, display, or redistribution so each use can be addressed explicitly.