Language and instructions
Check meaning across Arabic varieties, appropriate register, constraints, and corrections across multiple turns.
Compare model versions using the language, documents, and tasks your users rely on. License Arabic evaluation data or build a private suite around your product.
Arabic evaluation sets are also available to license with our smartphone speech dataset. Ask us about the available versions and counts.
Check meaning across Arabic varieties, appropriate register, constraints, and corrections across multiple turns.
Test evidence-supported answers, reading, extraction, tables, and cross-document questions. Identify retrieval misses separately from unsupported answers.
Measure turn-taking, corrected details, confirmation, and resolution. Executable tasks can check what actually changed.
Test both unsafe responses and unnecessary refusals, using explicit criteria and severity levels suited to your application.
Select an existing set or define tasks, Arabic varieties, sample sizes, and metrics for a custom evaluation.
Held-out items stay separate from training deliveries, with access and versioning defined for the project.
Custom reports can include failure examples, severity, uncertainty, reviewer disagreement, and changes between model versions.
LabelBench and our Arabic agent reliability methods beta describe specific studies; their scores are not blanket ratings for catalog products.
Both sound fluent. The evaluation shows which one captures a corrected address, asks for required confirmation, and completes the requested change.
Yes. A custom suite can use examples and policies you are authorized to share. We define expected behavior, tested language varieties, and scoring criteria with your team before running the evaluation.
Held-out evaluation items are kept separate from training deliveries. For a custom project, we define access, versions, and the separation of training and test material so the results remain useful.