What we learn building frontier-quality Arabic data — written for the teams shipping Arabic models into production. No fluff, no vendor theater: dialects, compliance and evaluation, treated with the seriousness they deserve.
An interactive, local-first crash test for what Arabic agents actually did—tool by tool, state by state.
78% on binary toxicity becomes 25.3% on the real taxonomy—and 22.4% across five dialect regions. The reproducible audit.
Collapse eats distribution tails first — and dialects are the tails. What the research means for Arabic, and the pipeline that avoids it.
Per-token pricing hides a 2–3× surcharge on Arabic. Where the tax comes from — and how to measure it on your own corpus.
MSA gets you reading. Dialect gets you understood. Where the real data gap sits — and what it costs in production.
What in-region really means under Saudi Arabia's PDPL, and the questions to ask any data partner before work begins.
Beyond accuracy: why translated benchmarks mislead, and a five-dimension protocol for measuring what matters in dialect.
Tell us the task. We'll scope a pilot that proves the quality on your own data.
Talk to us