Books and edited text
Choose Arabic books or professionally edited MSA organized by subject. The edited collection covers material through 2026.
License books, subject-organized MSA text, and structured question-answer data. Each collection has a clear training task, with its own counts and licensing details.
Choose a collection to see its counts, intended use, and licensing details.
For model teams that need longer Arabic texts for language modeling or deeper coverage of a particular subject.
View dataset 400,000 multiple-choice Q&A itemsFor teams training Arabic educational assistants or models that answer STEM questions.
View dataset 200,000 sentencesFor teams building Arabic entity extraction into search, document processing, or language-understanding systems.
View dataset 180,000 question-answer pairsFor teams building Arabic assistants whose users ask questions about the Arab region.
View dataset 40,000 question-answer pairsFor teams building English-language assistants that need to answer questions about the Arab region.
View dataset Ask about available volumeFor model teams adapting an Arabic model to a subject where edited language and clear content rights matter.
View datasetOur team can add data or annotation for your particular model. These additions are agreed separately.
Choose Arabic books or professionally edited MSA organized by subject. The edited collection covers material through 2026.
Arabic STEM multiple-choice questions sit alongside separate Arabic and English Q&A collections about the Arab region.
The 200,000-sentence MSA named-entity-recognition collection supports extraction from formal Arabic text.
Tell us which professional, educational, technical, or regional material you need. We can plan authorized sourcing and annotation around that need.
Review a sample, subject coverage, register, date coverage, and the unit used to report volume.
Custom preparation can include document boundaries, source records, publication dates, license identifiers, cleaning history, and deduplication records.
The edited MSA collection includes commercial AI-training rights in the agreement. Training, retrieval, display, and redistribution are separate licensed uses.
Start with edited material in the relevant subject, add questions tied to that material, and use a separate evaluation set to check the assistant’s answers.
They are separate products: 180,000 Arabic pairs and 40,000 English pairs. Parallel translations or item-level alignment are not assumed; request the collection specifications if your application needs matched pairs.
Commercial AI-training rights are written into the agreement. If you also need to display passages, use the material for retrieval, or redistribute it, those uses need to be specified separately.