Arabic RLHF & human feedback data
Preference rankings, rubric scores, rewrites and reward-model data for Arabic LLMs. Every judgment comes from a native rater matched to your target dialect and domain, and is calibrated against gold standards.
Arabic RLHF is the human-feedback stage of training an Arabic LLM. Native speakers compare, score and rewrite model responses, and a reward model or a preference optimizer learns from those judgments what a good Arabic answer is. Bayanat Labs supplies this data to teams building or adapting Arabic models: preference rankings, rubric scores, SFT rewrites, reward-model sets and red-team signals. Each judgment is made by a vetted native rater matched to your target dialect and domain.
Who rates matters, because a model is aligned to its raters. The InstructGPT authors report that their labelers were mostly English speakers in the US or Southeast Asia, and that they agreed with each other about 73% of the time. Arabic adds questions those raters cannot judge: MSA or dialect, which dialect, what register, and whether the answer fits the culture.
What we deliver
Pairwise & listwise ranking
Raters pick the better of two or more responses, mark how much better it is and tag reason codes. This is the core Arabic preference data for DPO and reward models.
Rubric scoring
Separate scores for accuracy, register, dialect, cultural fit and safety. Use them to train regression reward models or to find where a checkpoint fails.
SFT writing & rewriting
Rejected answers rewritten as gold responses in the target variety. For larger volumes, see Arabic LLM training data.
Reward model data
Balanced pairs, hard negatives that differ on one criterion only, and explicit tie and "both unacceptable" labels. A held-out set tests the reward model itself.
Safety & red-team signals
Adversarial prompts written in dialect, plus judgments on refusals and over-refusals. For full campaigns, see Arabic AI red teaming.
Domain-expert feedback
Physicians, lawyers, bankers and engineers check accuracy in their own fields, where a fluent but wrong answer does the most damage.
What "good" means in an Arabic response
Generic helpful-and-harmless rubrics miss the failures Arabic users notice first. Arabic speakers do not share a single culture, and models already lean toward Western entities in Arabic contexts. So each of these becomes a separately scored criterion. The rubric below is illustrative; we finalise it with you.
| Criterion | What the rater checks | Typical Arabic failure |
|---|---|---|
| Instruction following | Task, length and format are as requested | Ignores "keep it short" |
| Accuracy | Facts and procedures are correct for the user's country | Correct for Egypt, wrong for Saudi Arabia |
| Register | Formality matches the product's register policy | Stiff, translated-sounding MSA in reply to a casual question |
| Dialect match | The answer stays in the required variety | Egyptian عايز in a Gulf answer (أبغى) |
| Cultural fit | References, honorifics, dates and religious content suit the user | Western defaults, or Gregorian dates where Hijri is expected |
| Language quality | Grammar, hamza, taa marbuta and natural phrasing | Calques such as "your satisfaction is our top priority" |
| Safety | Follows the policy in both directions | Refuses a harmless medical question, or misses a harmful request phrased in dialect |
Raters score each criterion from 0 to 4, the same Likert range NVIDIA's HelpSteer2 used for its attributes. They then give an overall preference and its strength: much better, slightly better, tie, or both unacceptable.
Worked example: one prompt, two answers
Prompt (Saudi dialect): وش أحسن طريقة أرد فيها على عميل معصّب في الواتساب؟ أبي رد قصير. In English: "What's the best way to reply to an angry customer on WhatsApp? I want a short reply."
Response A: عزيزي العميل، نشكرك على تواصلك معنا ونأسف لأي إزعاج قد تكون تعرضت له. إننا نقدّر صبرك، ورضاك هو أولويتنا القصوى… This is formal MSA: "Dear customer, thank you for contacting us and we apologize for any inconvenience… we appreciate your patience, and your satisfaction is our top priority…" Seven generic tips follow.
Response B: «هلا [الاسم]، حقك علينا ونعتذر عن اللي صار. عطني رقم الطلب وأتابع موضوعك بنفسي اليوم.» وخلك هادي ولا تبرر كثير. In English: "'Hi [name], you're right to be upset and we're sorry for what happened. Send me the order number and I'll follow it up myself today.' Stay calm and don't over-explain."
B wins on three counts. It is short, as requested. It uses natural Gulf customer-care phrasing. And it gives a message ready to send, with a next step. A is grammatical and polite, and that is the trap: a rater without Gulf fluency will often reward it for sounding thorough. Reward models are already easily swayed by length. Add a formality bias and the model learns that stiff MSA is "good Arabic". A dialect-matched rater records the preference, how strong it is, and reason codes (ignored length, register mismatch, translationese), so every judgment can be audited later.
Who rates: dialect-matched natives and domain experts
We cover 25+ Arabic varieties, including MSA, Gulf (Najdi and Emirati among them), Egyptian, Levantine, Iraqi and Maghrebi. Raters are matched to two things: the variety of the prompt and the variety the product should answer in. A Levantine rater does not judge Najdi register. Specialist prompts go to licensed professionals and editors. Most RLHF work is text, but the same workflow can rank audio, image and video outputs. Dialect coverage explains why this sets your model's ceiling.
Contributors pass dialect and domain screening before they see client data. They then qualify on gold pairs that have already been adjudicated. Gold items stay mixed into live work, so drift shows up per rater instead of in your next training run.
How an Arabic RLHF project runs
- Scope (Vet). We agree dialects, register policy, domains, safety policy, training method (DPO, reward model or online RL) and output schema.
- Rubric and gold set (Train). We write the rubric with your team, build adjudicated gold pairs and qualify raters on them.
- Pilot (Produce). Raters work a fixed batch of your model's outputs. We analyse agreement, gold accuracy and the cases where raters disagree.
- QA and adjudication (QA). Some items are rated by more than one rater. Linguists adjudicate disagreements, and anything that fails gold or rubric checks is reworked.
- Scale and deliver (Deliver). Versioned releases in your schema, with rater metadata and an audit trail on every task.
How we measure quality and agreement
Each label type gets the agreement statistic that fits it. Pairwise choices get raw agreement and Cohen's or Fleiss' kappa. Ordinal rubric scores get Krippendorff's alpha. Each rater is also scored for accuracy on the gold items mixed into their work. When raters disagree, we separate two causes:
- Rubric errors, such as a missed factual mistake or a misread instruction. These are adjudicated and reworked.
- Legitimate disagreement, such as Gulf and Levantine raters differing on what sounds polite. We keep these labels and tag them with the rater's variety. Averaging them away would quietly align the model to one community.
A tip for buyers: be wary of near-perfect agreement on subjective pairs. In InstructGPT, training labelers agreed with each other 72.6% of the time and held-out labelers 77.3%. Much higher figures on open-ended Arabic prompts can mean raters are keying on surface cues such as length or formality rather than content.
Delivery formats
| Deliverable | Key fields | Typically trains |
|---|---|---|
| Preference rankings | prompt, chosen, rejected, strength, reason codes | DPO, reward models |
| Binary ratings | prompt, completion, label | KTO |
| Rubric scores | 0–4 per criterion, overall | Regression reward models, evals |
| SFT rewrites | prompt, gold completion, original | SFT |
| Red-team sets | prompt, attack type, harm category, safety rating | Safety tuning, evals |
Files arrive as JSONL or Parquet. We use the schemas that Hugging Face TRL trainers expect, or yours. Rows can carry rater ID, rater variety and rubric version, so you can filter or reweight by community.
When to automate and when to use human raters
Models are useful for generating candidate responses, dropping malformed outputs and pre-sorting obvious pairs. Final Arabic judgments are another matter. In our LabelBench audit, a local language model reached 78% exact-match accuracy on binary Arabic toxicity on the fixed held-out manifest. On the real seven-label taxonomy it fell to 25.3%, and on five dialect regions it reached 22.4% (chance is 20%), calling 75% of Gulf examples Egyptian. A judge that cannot tell dialects apart cannot score dialect match. Safety shows the same gap. The authors of the human-curated ASAS benchmark (801 prompts, seven models) report that most models fail to defend against 50% of unsafe prompts, and that automated safety judges such as GPT-4o perform poorly compared with human annotators.
So we move a judgment out of the human queue only when Safe Automation Coverage, measured on your data, shows it is safe. We also keep a human-labelled slice to audit any AI-feedback pipeline. The failure it catches is described in model collapse speaks MSA.
Security and data residency
Prompts and chat logs often contain customer data. Data is hosted in-region by default, with on-prem and private-cloud options. Contributors work under NDA with least-privilege access, and every task is audit-logged. See the PDPL data residency checklist.
Use cases by industry
Finance and banking: product answers in dialect that stay within your advice policy. Healthcare: physician-rated accuracy and safe escalation. Government: formal answers to citizens' dialect questions. Legal: lawyer-reviewed answers for the right jurisdiction. Telecom, media and retail: customer-care tone and brand voice for each market.
In-house, crowd platform or specialist vendor?
| In-house team | Crowd platform | Bayanat Labs | |
|---|---|---|---|
| Dialect-matched raters | Whoever you can hire | Often self-reported | Screened per variety |
| Rubric and gold design | Yours to build | Usually yours | Written with you |
| Adjudication | Your reviewers' time | Often majority vote | Built into the workflow |
| Domain experts | Costly to add | Rare | Licensed professionals |
| Residency | Full control | Often a global workforce | In-region by default |
Whichever you choose, ask each option four things. Which variety does each rater speak? How was the rubric validated? What are overlap agreement and gold accuracy per rater? Where do the raters sit? Then check that the data moved the model, using our protocol for evaluating Arabic LLMs or an independent Arabic LLM evaluation.
Start with a pilot
A fixed-scope pilot runs on your own model's outputs. It includes a co-written rubric, gold-standard QA, and a quality report covering agreement, gold accuracy and disagreement patterns, usually within two weeks. After the pilot, we scale as a managed, embedded or enterprise engagement.
Frequently asked questions
How is Arabic RLHF different from English RLHF?
The mechanics are the same: rank, score, rewrite, train. What changes is the rubric and who rates. An Arabic answer can have the right content in the wrong variety. It might use MSA when the user wrote in dialect, use Egyptian forms in a Gulf product, or fall back on a Western cultural default. Non-native raters miss these errors, so you need dialect-matched raters and explicit criteria for each one.
Can we translate English preference data into Arabic?
Translation is fine for seeding prompts, but it does not give you Arabic preferences. A translated pair still carries the norms of its original raters, plus translated phrasing that a reward model can learn to favour. If you start from translated data, have native raters re-rank it and rewrite it in the target variety. Then check the result against a natively written evaluation set.
Should an Arabic model answer in MSA or in dialect?
It depends on the product, so we write it down as a register policy instead of leaving it to each rater's taste. A banking or government assistant may answer dialect questions in clear, light MSA. A consumer chatbot may reply in the user's dialect. Without a written policy, raters split by personal preference and the reward model learns noise.
How many Arabic preference pairs do we need?
No single number fits every project. It depends on the training method, the size of the behaviour gap and how consistently raters apply the rubric. Size it by measurement: run the pilot, train on growing subsets, and keep collecting while evaluation gains are still rising. A narrow fix, such as tone in one domain, needs far less data than a general-purpose reward model.
Do you rate our model's outputs or supply the responses?
Either. The most useful signal usually comes from rating samples from your current checkpoint, because that is the behaviour you want to change. We can also write prompts in the target dialect, produce gold rewrites, or rank candidate responses from several models. DPO skips the separate reward model, but it still trains on these same chosen-and-rejected pairs.