Arabic retrieval starter kit
Twelve passages, twelve questions and thirty graded judgments that show where Arabic enterprise search and RAG go wrong: dialect questions, English terms inside Arabic, evidence in the other language, outdated policy versions, rules for a different customer, and questions the sources cannot answer.
Enterprise search fails differently from public benchmarks. As a 2025 study of Arabic–English RAG over corporate documents points out, earlier work mostly relied on benchmarks built from open-domain sources, especially Wikipedia. Inside a company, employees ask in Gulf Arabic about a policy written in formal Arabic, mix in English terms, or ask in Arabic about a document that only exists in English. The index holds last year’s policy next to this year’s, and a rule for business customers that looks almost identical to the one for individuals. The same study found that retrieval was the main bottleneck, with substantial performance drops when the question and the supporting document are in different languages.
This kit packs those cases into a fixture small enough to read in ten minutes and run in one. It is fully synthetic: Bayanat Labs wrote every passage, question and judgment for a fictional telecom company. Use it to test that your pipeline, your scoring and your answer checks behave as you expect before you spend money on a full evaluation.
بالعربية: مجموعة بيانات صغيرة ومفتوحة لتقييم البحث والتوليد المعزز بالاسترجاع (RAG) بالعربية والإنجليزية: أسئلة باللهجة الخليجية والفصحى، ومصطلحات إنجليزية داخل الجملة العربية، وأدلة بلغة غير لغة السؤال، ونسخ قديمة من السياسات، وأسئلة لا تجيب عنها المصادر. جميع النصوص اصطناعية كتبتها Bayanat Labs لشركة خيالية. لبناء مجموعة كاملة من مستنداتك، راجع بيانات البحث والاسترجاع بالعربية.
What is in the kit
| File | Contents |
|---|---|
| corpus.jsonl | 12 passages in Arabic and English, each with a version, effective date, status (current or superseded) and customer group. Two are translations of Arabic passages; one is a superseded version. |
| queries.jsonl | 12 questions: Gulf colloquial, MSA, Arabic with English terms, and English. Each has a note explaining what it tests. |
| qrels.txt | 30 graded judgments in TREC qrels format (query_id 0 doc_id grade). |
| judgments.jsonl | The same judgments with a reason code, such as hard_negative_superseded_version. |
| answers.jsonl | The expected grounded answer, exact evidence quotes, whether the question is answerable, the expected behavior when it is not, and the passages an answer must not rely on. |
| SHA256SUMS.txt, LICENSE.txt | Checksums for every file, and the CC BY 4.0 license. |
Relevance grades follow one rule: 3 the passage answers the question; 2 it answers part of it; 1 it is on topic but does not answer; 0 it is not relevant, including hard negatives that look right but are not.
Six failure cases it isolates
| Case | Example in the kit | What a good system does |
|---|---|---|
| Dialect question, formal source | Q01 أبي أرجّع الجوال اللي اشتريته قبل أسبوع، أقدر؟ | Retrieves the MSA returns policy and answers from the current version. |
| Arabic question, English evidence | Q03 وش مدة الـ warranty لو الجهاز فيه عيب مصنعي؟, Q10 | Finds the English warranty and call-recording passages and answers in Arabic. |
| English question, Arabic evidence | Q04 “Can a business customer return devices after three weeks?” | Ranks the Arabic business-customer policy above the English individual policy. |
| Superseded version | Q02, Q12 | Treats the 2025 policy as a hard negative: it allowed opened returns with a 10% fee and refunds as account credit. |
| Different customer group | Q01, Q04 | Keeps the 30-day business rule out of answers for individuals, and the 14-day rule out of answers for businesses. |
| Partly answerable or unanswerable | Q07, Q08, Q09 | Says what the sources cover and what they do not. It does not invent a porting fee or apply the returns rule to a device bought elsewhere. |
A worked example
Q01 asks, in Gulf Arabic, whether a phone bought a week ago can be returned. Five passages are judged:
| Passage | Grade | Why |
|---|---|---|
| RET-AR-2026, current returns policy for individuals (Arabic) | 3 | States the 14-day window, packaging and invoice conditions. |
| RET-EN-2026, the same policy in English | 3 | The same evidence in the other language. |
| RET-AR-FAQ, returns FAQ | 1 | On topic, but points to the policy without giving the rule. |
| RET-AR-2025, the superseded policy | 0 | Hard negative: a 7-day window that no longer applies. |
| RET-AR-BIZ-2026, policy for business customers | 0 | Hard negative: a 30-day window for a different customer group. |
The expected answer is yes, as long as 14 days have not passed since delivery and the phone is in its original packaging with the invoice; an opened phone is accepted only with a manufacturing defect. An answer that cites the 7-day rule or the 30-day rule fails, even if the retriever also found the right passage.
How to score it
- Retrieval. Index
corpus.jsonl, run the 12 questions and write a TREC run file. Score it againstqrels.txtwith trec_eval, pytrec_eval or ir_measures, for example nDCG@10, and Recall@5 counting only grades 2 and 3 (-l 2in trec_eval). Unjudged passages count as not relevant. - Answers. Generate answers from your top passages and check each against
answers.jsonl: does the answer agree with the evidence quote, does it avoid themust_not_usepassages, and does it decline or ask for detail on Q08 and Q09? - Report them separately. A system can retrieve the right passage and still answer from the old version. One combined score hides which step failed.
What this kit is not
- Not a benchmark. Twelve questions are enough to test a pipeline and a scoring setup, not to rank systems.
- Not real data. The company, policies and numbers are fictional. There is no customer or personal data in it.
- Not a sample of a Bayanat dataset. It shows the format and the cases; commercial retrieval datasets are built to order on a client’s authorized documents.
- Single-author judgments. The grades were written with the fixture, so no reviewer agreement is reported. A real project measures and reports disagreement between reviewers.
Build the real version on your documents
Bayanat Search builds this kind of dataset around your own authorized corpus: questions in the words your users actually use, graded relevance from native Arabic reviewers, hard negatives drawn from your outdated and look-alike documents, unanswerable questions, and reviewer reasoning, including where reviewers disagree. Pair it with Bayanat Bench to keep a held-out test set separate from training data, and with Bayanat Documents when your evidence sits inside scanned forms and contracts.
Cite or reuse
Free to use, adapt and redistribute under CC BY 4.0 with attribution:
Bayanat Labs (2026). Arabic Retrieval Starter Kit v0.1. https://www.bayanatlabs.com/research/arabic-retrieval-starter-kit
Found an error, or want a larger version for a specific domain? Write to hello@bayanatlabs.com.