Arabic Preference Data for RLHF & DPO: How to Collect It Natively
Many public Arabic preference datasets are translated or model-generated, so they risk teaching an Arabic model English preferences in Arabic words. Here is how to collect preference pairs natively: the rubric, the raters, the agreement checks and the file formats your trainer expects.
Arabic preference data is a set of human judgments on model answers to Arabic prompts: which of two or more responses is better, by how much, and why. Reward models, DPO and related methods learn what a good Arabic answer is from these judgments. To be useful, the data has to be collected natively: prompts written in the varieties your users actually write in, and responses judged by raters who speak those varieties, against a rubric that scores dialect, register and cultural fit alongside helpfulness and safety.
Many public Arabic preference datasets are translated or model-generated, with thin documentation, and generic RLHF guides never say what an Arabic rubric should contain. This guide covers why translation breaks the signal, a rubric you can adapt, a worked Najdi example, collection methods, rater matching and calibration, and the formats DPO and reward-model trainers expect.
What is Arabic preference data?
A preference record has a prompt, two or more candidate responses, and a judgment: a pairwise choice, a ranking, a rating on a scale, or a binary acceptable/unacceptable label. Llama 2's annotators, for example, chose between two responses and then rated how strongly they preferred it: significantly better, better, slightly better, or negligibly better/unsure. They judged helpfulness and safety separately.
What makes the data Arabic is not the script. The prompts must reflect how Arabic speakers actually ask, in MSA, dialect or a mix; the responses must cover the choices an Arabic model faces, such as which variety to answer in; and the raters must be able to tell a natural answer from a translated one. Miss any of these and the model learns its data-maker's preferences, written in Arabic.
What public Arabic preference datasets exist?
A few Arabic preference sets are public on the Hugging Face Hub; read the cards before training. As of September 2026, the most visible ones:
| Dataset | How the pairs were made (per its card or paper) | Use it for |
|---|---|---|
| FreedomIntelligence/Arabic-preference-data-RLHF | About 11.5k rows of instruction, chosen and rejected. The repo has no dataset card, so source, method and licence are undocumented. | Exploration only, until provenance is known |
| 2A2I/argilla-dpo-mix-7k-arabic | 7,500 rows with chosen and rejected ratings, tagged synthetic. The name points to Argilla's English dpo-mix-7k, and the card does not document translation or review. | Pipeline smoke tests, warm-start experiments |
| 2A2I/NoRobots-SambaLingo.Arabic.Chat-DPO | 9.5k rows. Chosen responses are translated human answers from an Arabic version of No Robots; rejected responses are AI-generated. The card text says CC BY-NC 4.0, while the repo metadata says Apache 2.0; check before use. | Research, noting that chosen and rejected differ in origin |
| Juhaina alignment data (CamelEval paper) | The authors report about 40,000 alignment rows, with native Arabic speakers giving preference feedback on pairs of model answers, used for ORPO training. | A reference design; the paper describes the process |
The pattern: the open sets in the table are translated, synthetic or undocumented, and the one native human collection is described in a paper, not released with a dataset card. Public sets are fine for testing a pipeline, not for defining what your model should prefer.
Why do translated preference pairs fail?
Translating an English preference set looks efficient, since the judgments already exist. But a judgment belongs to the original text, not the translation.
- Translationese becomes the signal. When Cohere For AI researchers built multilingual preference data covering 23 languages, Arabic among them, they avoided simply translating English pairs. They warned that translation artifacts can hurt performance and that re-translating the same pairs reduces diversity. In their setup, translated completions were ranked as the "bad" completion 91% of the time, so the preference label largely acted as a proxy for "this was translated". Any dataset where translation correlates with one side can teach that shortcut.
- The content stays foreign. Names, holidays, institutions and assumptions come from the source language. When the CIDAR team machine-translated 9,109 English instruction pairs (from AlpaGasus) into Arabic, reviewers modified 64.5% of them, for translation errors or Western-specific content, before they fit an Arabic dataset.
- The dialect axis never appears. Translation output defaults to MSA, so a translated set rarely contains the judgment Arabic users notice first: whether a dialect question should get a dialect or an MSA answer.
- Evaluation inherits the gap. M-RewardBench machine-translated RewardBench into 23 languages, including Arabic, using Google Translate. All reward models tested scored lower on average than on the English original, and scores rose with translation quality. The authors also found that reward-model preferences can change substantially from one language to another.
Translated prompts can still seed coverage, if native raters rewrite them and fresh responses are judged in Arabic. The same applies to model-generated data: as our piece on synthetic Arabic data and model collapse explains, a generator's MSA skew compounds when its output becomes training data.
An Arabic preference rubric
A rubric turns "which answer is better?" into questions a rater can answer consistently. For Arabic, split the criteria into gates, which a response must pass before anything else counts, and graded criteria, which decide between two acceptable answers. The rubric below is a template, not a standard; adapt the scales and the register policy to your product.
| Criterion | Question the rater answers | Type | Arabic failure it catches |
|---|---|---|---|
| Safety | Does the response follow the safety policy, in both directions? | Gate | Missing a harmful request phrased in dialect; refusing a harmless religious or medical question |
| Accuracy | Are facts, dates, procedures and details from the prompt preserved and correct for the user's country? | Gate | Correct for one Arab country, wrong for another; a time reference silently changed |
| Instruction following | Does it do the task at the requested length, format and tone? | Graded 0–4 | Ignores "keep it short" or "not formal" |
| Dialect fit | Is it in the variety your policy requires for this prompt? | Graded 0–4 | Egyptian forms in a Najdi conversation; MSA where the policy says to mirror the user |
| Register | Is the formality right for the product and the user? | Graded 0–4 | Letter-style MSA in reply to a casual message |
| Cultural fit | Do the references, greetings, honorifics, calendar and religious framing suit the user? | Graded 0–4 | Western defaults; an invitation that ignores prayer times |
| Language quality | Is it grammatical and idiomatic, without calques? | Graded 0–4 | Translated set phrases, hamza and taa marbuta errors |
Two additions make it workable. First, a register policy: for each product surface, state whether the model answers in MSA, the user's dialect, or light MSA with dialect vocabulary. Without one, raters split on taste and the reward model learns noise. Second, two extra outcomes: tie and both unacceptable. Forcing a choice between two bad answers teaches the model that the less bad one is good. Our free Arabic annotation guidelines template gives you a structure for documenting decisions like these.
A worked example: three answers to a Najdi prompt
The example below is illustrative: it was written for this guide, and the scores are not measured results. It shows how a native rubric separates answers that look alike in translation.
Prompt (Najdi): أبي رسالة قصيرة أعزم فيها الشباب في الدوام على عشاء عندي في البيت ليلة الجمعة بعد العشا. خلها ودية وما تكون رسمية. In English: "I want a short message inviting the guys at work to dinner at my place on laylat al-jumʿa after Isha. Make it friendly, not formal."
ليلة الجمعة (laylat al-jumʿa) is the night before Friday, which is Thursday evening. In common Arabic usage, a night belongs to the day that follows it. A literal English rendering, "Friday night", moves the dinner by a day.
Response A: الزملاء الأعزاء، يسعدني دعوتكم لحضور عشاء في منزلي مساء يوم الجمعة. أرجو تأكيد حضوركم في أقرب وقت ممكن. مع أطيب التحيات. In English: "Dear colleagues, I am pleased to invite you to a dinner at my home on Friday evening. Please confirm your attendance at your earliest convenience. Best regards."
Response B: يا هلا بالشباب، ودي أشوفكم عندي على العشا ليلة الجمعة بعد صلاة العشا. البيت بيتكم، وعطوني خبر عشان أرتب. In English: "Hey guys, I'd love to have you over for dinner on Thursday night after the Isha prayer. My home is your home. Let me know so I can get things ready."
Response C: يا جماعة، عايز أعزمكم على العشا عندي يوم الخميس بالليل بعد العشا. ياريت تيجوا! In English: "Guys, I want to invite you to dinner at my place Thursday night after Isha. Hope you can come!"
| Criterion | A | B | C |
|---|---|---|---|
| Safety (gate) | Pass | Pass | Pass |
| Accuracy (gate) | Fail: Friday evening, and "after Isha" is dropped | Pass | Pass |
| Instruction following | 1: short, but formal | 4 | 3 |
| Dialect fit | 0: formal MSA | 4: natural Najdi | 1: Egyptian عايز, ياريت تيجوا |
| Register | 0: letter style | 4 | 3 |
| Cultural fit | 2 | 4: hospitality phrase, prayer named | 3 |
The native ranking is B > C > A. C gets facts and tone right but speaks the wrong dialect, a graded miss. A fails a gate: a guest who reads it arrives a day late. Written as pairs, that gives B over A (significantly better), B over C (better) and C over A (better).
Now picture a non-Arabic-speaking reviewer working from back-translations. A comes back as "Friday evening", matching a translated prompt that reads "Friday night", so the error vanishes, and A reads as the most polished. That reviewer could plausibly rank A first; thousands of such judgments teach a reward model that formal MSA with quietly changed details is the best Arabic. Note also why بعد صلاة العشا in B is better than a bare بعد العشا. In speech, العشا means both dinner and the Isha prayer, so naming the prayer removes the ambiguity. Only a fluent rater credits that.
Pairwise, ranking or Likert: which should you use?
The collection method decides which training methods you can use and what each judgment costs in rater effort:
| Method | The rater… | Yields | Best for | Watch for |
|---|---|---|---|---|
| Pairwise with strength | Picks the better of two and says by how much | One pair plus a margin | DPO, Bradley-Terry reward models | Forced choices between two bad answers; allow ties |
| Ranking (K responses) | Orders four or more responses | K(K−1)/2 pairs per prompt | Efficient reward-model data | Pairs from one ranking are correlated |
| Likert per criterion | Scores each response on each rubric row | Absolute scores | Regression reward models, diagnostics, evals | Scale drift between raters; anchor every point |
| Binary acceptable | Marks one response good or bad | Unpaired labels | KTO-style training | Loses the "slightly better" signal |
The evidence favours mixing methods. InstructGPT's labelers ranked 4 to 9 responses per prompt; because comparisons from one ranking are highly correlated, shuffling them into one dataset made the reward model overfit, so the authors trained on each prompt's comparisons as a single batch element. Llama 2 used preference strength to set a larger loss margin for clearly different pairs. NVIDIA's HelpSteer2-Preference compared pairwise (Bradley-Terry) and rating-based (regression) reward models on matched data and got its best model by combining the two.
A practical Arabic default: pairwise preferences with a strength label for training, plus 0–4 scores on dialect, register and cultural fit for diagnostics. The scores explain why a checkpoint loses ("accurate but always MSA"), which one preference bit cannot.
How do you match raters to dialect and domain?
A rater can only judge dialect fit in a variety they speak natively, and the prompt's variety and the answer's variety can differ:
- Prompt and answer in the same dialect (a Najdi user, with a policy of mirroring the user): the rater must be a native speaker. A Levantine rater can judge helpfulness, not whether the Najdi sounds natural.
- Dialect prompt, MSA answer (a government or banking assistant): the rater must fully understand the dialect and judge MSA quality. Misreading the question is the common failure.
- Arabizi or mixed-script prompts: raters must know the users' written habits, not just their speech.
- Specialist prompts (medical, legal, financial, religious): a domain professional judges accuracy and a dialect-matched rater judges language.
Record each rater's variety and domain on every row, so you can see whether preferences shift by community. The research expects them to: PRISM, built from 1,500 participants in 75 countries, was designed around cross-cultural disagreement on value-laden prompts. For why variety coverage sets your model's ceiling, see MSA vs dialectal Arabic.
How do you measure rater agreement and calibration?
Agreement shows whether the rubric is applied consistently, not that the labels are right. Build in three controls:
- Gold pairs. Adjudicated pairs with a written rationale, used to qualify raters, then mixed invisibly into live work. Per-rater gold accuracy is your earliest drift signal.
- Overlap. A fixed share of items rated by two or more raters. Report raw agreement and a chance-corrected statistic: Cohen's or Fleiss' kappa for pairwise choices, Krippendorff's alpha for ordinal rubric scores.
- Calibration sessions. Regular reviews of disagreements that end in a rubric clarification or new gold item, not a majority vote.
Set realistic expectations. On its instruction-following prompts, InstructGPT reports that its training labelers agreed with each other 72.6 ± 1.5% of the time, and held-out labelers 77.3 ± 1.3%. Near-perfect agreement on subjective Arabic pairs is a warning sign: raters may be keying on length or formality.
When raters disagree, separate the cause. Rubric errors, such as a missed gate, are adjudicated and reworked. Legitimate disagreement, such as Gulf and Levantine raters differing on what sounds polite, is kept and tagged by variety; averaging it away aligns the model to your rater pool's majority.
Automated judges need the same scrutiny. In our LabelBench audit, a local language model reached 78% exact-match accuracy on binary Arabic toxicity on the fixed held-out manifest. It fell to 25.3% on the real multi-label taxonomy, and 22.4% across five dialect regions (chance is 20%). A grader that cannot identify the dialect cannot score dialect fit. After training, check the gains on a natively written test set, using our Arabic LLM evaluation protocol.
What format do DPO and reward models need?
Hugging Face's TRL documentation (as of September 2026) defines the shapes its trainers expect. A preference row has prompt, chosen and rejected fields, as strings or chat-style message lists. DPOTrainer and ORPOTrainer recommend an explicit prompt; RewardTrainer recommends the implicit form, with the prompt inside both completions. KTOTrainer also accepts unpaired rows: one completion and a true/false label. The docs warn that unpairing is only safe if every chosen response is good and every rejected one bad, which is what a "both unacceptable" label tells you.
Keep trainer columns clean and carry the Arabic-specific information as metadata:
| Field | Example value | Why keep it |
|---|---|---|
| prompt_variety | najdi | Slice reward-model accuracy by dialect |
| target_variety | mirror_user / msa | Links each judgment to the register policy |
| strength | significantly_better | Margin in the loss, or filtering out weak pairs |
| reason_codes | accuracy_gate, dialect_mismatch | Audit why the chosen response won |
| criterion_scores | {"dialect_fit": 4, "register": 4} | Regression reward models and diagnostics |
| rater_id, rater_variety | r_117, najdi | Agreement, drift and community-level analysis |
| response_source | ckpt_0915 / human_rewrite | Stops an origin artifact from becoming the signal |
Two hygiene rules. Normalize chosen and rejected text identically (diacritics, tatweel, alef and yaa forms), or the model learns the formatting difference instead of the quality difference. And avoid origin splits: if every chosen answer is a human rewrite and every rejected one is model output, the pair teaches "sound human", the same shortcut the translation study found.
How Bayanat approaches Arabic preference data
Bayanat Labs is a Riyadh-based Arabic data company, and preference ranking is part of our human feedback work. Contributors are vetted native speakers, linguists and licensed domain professionals who pass dialect and domain screening before touching client data. Work is calibrated against gold standards, disagreements go through adjudication and rework loops, and every task has an audit trail. For preference ranking, rubric scores, SFT rewrites and reward-model sets, see our Arabic RLHF and human feedback service. When the goal is safety behaviour, adversarial prompts written in dialect come from our Arabic red teaming work.
Checklist: collecting Arabic preference data
- Write prompts natively in each target variety; rewrite translated prompts, never use them as-is.
- Publish a register policy per product surface before rating starts.
- Sample both responses from your checkpoint or similar-strength models, so pairs differ in quality, not origin.
- Use gates (safety, accuracy) and graded rows (instructions, dialect, register, culture, language).
- Allow "tie" and "both unacceptable"; route unacceptable pairs to native rewrites.
- Match raters to prompt and answer variety; add domain experts for specialist prompts.
- Qualify raters on adjudicated gold and keep gold mixed into live work.
- Report overlap agreement per label type, and question suspiciously high agreement.
- Keep legitimate disagreement, tagged by rater variety; deliver TRL rows with that metadata.
- Normalize text identically, then validate on a natively written test set, by dialect.
- Translation moves words, not judgments. Translated pairs keep their original raters' preferences, and translation artifacts can become the signal.
- The visible open Arabic preference sets are translated, synthetic or undocumented. Use it to test your pipeline, not to define what your model should prefer.
- Score dialect, register and cultural fit explicitly. Make safety and accuracy gates, and write a register policy so raters do not split on taste.
- Pairwise with strength, plus rubric scores, is a strong default. The pairs train the model; the scores explain its failures.
- Who rates is part of the dataset. Match raters by variety and domain, calibrate on gold, tag every row.
Frequently asked questions
Can I use an LLM instead of human raters to label Arabic preference pairs?
For pre-sorting, yes: a model can drop malformed outputs and flag obvious pairs. For final labels, only after measuring it against native raters on your own data, because automated judges struggle with the dialect and register axes that matter most in Arabic. Our guide to LLM-as-a-judge in Arabic covers where automated graders break and how to audit them.
Should the chosen response be written by a human?
Usually not for DPO. If every chosen answer is human-written and every rejected answer is model-written, the model can learn to tell human style from model style instead of good from bad. Sample both responses from your own checkpoint, or from models of similar strength, so they differ in quality rather than origin. Put human gold rewrites into your SFT set, or label them as a separate source.
Do we need separate preference data for each Arabic dialect?
Not necessarily a separate dataset, but you need coverage of every variety your users write in and a policy for which variety the model answers in. That means prompts written natively in each variety, matched raters, and a variety tag on every row so you can check the reward model per dialect. Pooled scores hide a weak dialect behind a strong MSA average.
How should ties and 'both bad' pairs be handled in DPO?
DPO needs a strict winner, so ties are dropped from the DPO file. Keep them anyway: they are useful for testing the reward model and for spotting rubric gaps. When both responses are unacceptable, neither belongs on the chosen side. Send the prompt for a native rewrite, which then becomes an SFT example or, clearly labelled, a new chosen response.
What is the difference between preference data and SFT data?
SFT data gives the model one good answer per prompt to imitate. Preference data records a judgment between answers: which is better, and often by how much. SFT teaches a format and a voice; preference data teaches trade-offs, such as concise versus thorough or dialect versus MSA, that a single demonstration cannot capture.