Arabic LLM evaluation services for models and agents
Independent evaluation for teams choosing, fine-tuning or shipping Arabic models: native raters for each variety, private held-out sets, state-based agent checks and a report you can defend.
Bayanat Labs provides Arabic LLM evaluation services: independent, blind, dialect-stratified human evaluation of models and agents by vetted native speakers and licensed domain professionals. You learn how a model performs on the varieties, domains and tasks you will actually deploy, reported per dialect and per dimension, with every raw judgment attached. We are model-agnostic and never compete with the models we evaluate.
What our Arabic model evaluation covers
Model selection
Blind side-by-side comparison of candidate models, closed or open-weight, on prompts drawn from your use case.
Private custom benchmarks
Held-out prompts authored natively in each target variety, with written rubrics and gold-scored anchors. Never published.
Fine-tune regression testing
One frozen set re-run on every checkpoint, so an MSA gain cannot hide a Gulf regression.
Preference and rubric scoring
Pairwise preference, rubric scores and error tags from raters who speak the variety they judge.
Safety and cultural alignment
Religious, social-norm and regional sensitivity, including over-refusal. Adversarial work runs as Arabic AI red teaming.
Agent evaluation
State-based checks on tool calls and final system state, not on what the reply claimed.
The five dimensions we score
Each is reported per variety and domain, never blended into one "Arabic score". The rationale is in our methodology guide, how to benchmark an Arabic LLM you can trust.
| Dimension | What we check | Example probe (illustrative) |
|---|---|---|
| Dialect comprehension | Dialect, code-switching and Arabizi input read correctly | One intent, three varieties: أبي أغيّر موعدي (Gulf), عايز أغيّر ميعادي (Egyptian), بغيت نبدّل الموعد ديالي (Moroccan), all "I want to change my appointment". Read as MSA, أبي means "my father". |
| Generation register | The variety and formality your product specifies, without translationese | A dialect question where the style guide requires a dialect reply. LLMs understand dialect better than they produce it, largely because they are reluctant to generate it (AL-QASIDA). |
| Regional factuality | Local institutions, regulation, places and calendars | "Which authority supervises personal data protection in Saudi Arabia?" (SDAIA) |
| Cultural and safety alignment | Norm sensitivity; over- and under-refusal | A benign question about religious practice that should be answered, not refused |
| Domain tasks | Your real jobs, judged by licensed professionals | A physician grades a triage reply to symptoms described in Levantine dialect |
Why public Arabic benchmarks aren't enough for a deployment decision
Leaderboards are for shortlisting. They cannot tell you whether a model is ready for your customers.
- They mostly test recognition, not writing. The Open Arabic LLM Leaderboard v2 is mostly multiple-choice, with one generative task (ALRAGE) graded by an LLM judge, and most of the seven benchmarks in HELM Arabic are multiple-choice or exam QA. Users judge what the model writes.
- Translation leaves marks. OALL's maintainers dropped machine-translated tasks from v2 "due to inherent lower quality and possible cultural bias".
- Dialect is rarely broken out, and when it is, scores fall. Across 19 open-weight models (1B to 13B parameters), DialectalArabicMMLU reports average accuracy of 47.7% on five dialects, against 51.9% on MSA and 62.8% on English.
- Public items can leak into training data, and none reflects your traffic mix, domain or policy.
How an Arabic LLM evaluation runs
- Scope the decision. Selection, launch sign-off or regression; target varieties and their traffic share; domains; pass thresholds. Raters pass dialect and domain screening before seeing your data.
- Build rubrics, gold set and held-out prompts. Prompts are authored natively per variety, never translated from English. Raters calibrate on gold-scored anchors.
- Pilot a slice. A small run tests rubric clarity and agreement; ambiguous items are rewritten before scale.
- Judge with QA and adjudication. Model identities hidden, output order randomized, hidden gold items seeded to catch drift, an overlap subset scored by several raters.
- Report and hand over. Per-dialect profile, error taxonomy and raw judgments, plus a frozen set for the next checkpoint.
How we keep human judgments trustworthy
Inter-annotator agreement is measured on the overlap subset, per dimension and dialect, and reported with the results. Low agreement triggers a rubric fix, not a model verdict: the rubric is then measuring the raters. Disagreements are adjudicated, rejected work is reworked, and every judgment carries an audit trail. Blind, randomized presentation counters the position and verbosity biases documented in LLM judges (Zheng et al., 2023).
Can an LLM judge replace native Arabic raters?
Automate what is checkable: multiple-choice answers, output format, tool calls, final state. Keep native raters where judgment is the measurement: register, cultural fit, dialect fidelity, domain correctness. In a small 2026 study comparing five frontier LLM judges with native-speaker expert grades on Egyptian and Iraqi Arabic prompts, four of the five judges were systematically lenient, and implicit cultural reasoning was the main failure mode for automated grading (Abdoli et al.). Our LabelBench audit points the same way: on the fixed held-out manifest, a local language model scored 78% on binary Arabic toxicity, but exact-match accuracy fell to 25.3% on the real seven-label taxonomy, and it reached 22.4% on five dialect regions (chance is 20%). The workable hybrid: calibrate an automated judge on a human-scored slice of your data, route disagreements to people, and measure what can safely leave the human queue.
Arabic agent evaluation: check the state, not the sentence
A fluent answer is not a completed transaction. For agents we grade the tool trail and final system state against a declared goal, as τ-bench does by comparing the final database state with an annotated goal state. Arabic raises the stakes: tool-calling accuracy drops by an average of 5–10% when users interact in Arabic, whatever the language of the tool descriptions (Arabic Prompts with English Tools).
Our Arabic Agent Reliability Lab, a methods beta, shows why. In one receipt, an agent replied تم إلغاء الطلب ORD-SBX-1005 بنجاح ("Order ORD-SBX-1005 was cancelled successfully") when its only tool call was an order lookup, and the order stayed packed. For one disclosed system (Qwen3.5 9B with least-privilege tools), operational completion was 15 of 21 synthetic workflow families but communication quality was 5 of 21, so we report the two separately. These are development results on an open synthetic suite, not accuracy scores, rankings or a certification. A private audit applies the method to your workflows, tools and access modes, with repeated trials and dialect prompts written and reviewed by native speakers.
Deliverables
| Deliverable | Contents | Use it for |
|---|---|---|
| Evaluation report | Method, sample design, results per dimension with confidence intervals, limitations | Selection memo or launch sign-off |
| Per-dialect profile | Variety × dimension scores; where the model misreads dialect or defaults to MSA | Market-by-market go/no-go |
| Error taxonomy | Categorized failures with Arabic examples and severity | Fix list for prompts, guardrails or RLHF and preference data |
| Raw judgments | Every score with pseudonymous rater ID, rubric version and adjudication notes | Re-analysis and audit |
| Frozen held-out set | Prompts, rubrics and gold anchors | Regression reruns on each fine-tune |
| Agent receipts | Request, tool trail and before/after state per case | Engineering fixes |
Coverage, security and industries
Raters cover 25+ Arabic varieties, from MSA to Gulf, Egyptian, Levantine, Iraqi and Maghrebi dialects (why that must be named variety by variety: dialect coverage), with physicians, lawyers, bankers, engineers and editors for domain judgments. The same protocol applies to speech transcription and to Arabic-script OCR and document understanding. Data is hosted in-region by default, with on-prem and private-cloud options, least-privilege access, audit logs and contributors under NDA (see PDPL and data residency). We serve finance and banking, healthcare, government, legal, telecom, and media and retail.
In-house, crowd platform or specialist partner?
| Question | In-house team | Crowd platform | Specialist partner |
|---|---|---|---|
| Native raters for every target variety? | Only the varieties your staff speak | Often self-reported | Screened per variety |
| Licensed domain judges? | Costly to pull from their day jobs | Rare | Recruited per project |
| Independent of the model? | Grading your own work | Yes | Yes, if they build no models |
| Rubrics, agreement, adjudication? | You build them | You build them | Part of the service |
| Where is the data processed? | Under your control | Wherever the workers are | Ask; ours is hosted in-region by default |
Start with a pilot
A pilot is fixed-scope: one decision, two or three varieties, one domain. It includes gold-standard QA and a quality report, usually within two weeks. Then scale as a managed, embedded or enterprise engagement. Tell us what you need to decide.
Frequently asked questions
How much do Arabic LLM evaluation services cost?
We quote after scoping instead of publishing a rate card. Cost follows variables you control: target varieties, prompts per variety, raters per item, and whether judges must be licensed professionals; agent audits add sandbox setup. A fixed-scope pilot gives you a quality report and a cost basis for scaling, usually within two weeks.
How many prompts does an Arabic model evaluation need?
Enough per decision cell (each variety and dimension) that the confidence interval is narrower than the difference you care about. Illustratively, 200 items at a 70% pass rate give a 95% interval of about ±6 points; 50 items give about ±13. Size per dialect, not in total, or a thin Levantine sample decides a Levantine launch.
Can you evaluate a model we did not build, or a competitor's API?
Yes. We do data work only and are model-agnostic, so we have no model of our own to favor. Closed APIs, open-weight models and your own fine-tunes are compared side by side, with identities hidden from raters. Private endpoints can be evaluated on-prem or in your private cloud rather than exporting outputs.
What is the difference between LLM evaluation and red teaming?
Evaluation measures how a model performs on the traffic you expect: representative prompts scored against rubrics and reported as rates. Red teaming hunts for worst cases, such as dialect and Arabizi jailbreaks or harmful advice in a domain context, and reports findings by severity. Most launches need both; see Arabic AI red teaming.