Service · Training data

Arabic LLM training data: SFT, instruction and fine-tuning data

Instruction–response pairs, multi-turn dialect conversations, domain Q&A and rewrites, written in Arabic by vetted native speakers and domain professionals rather than translated from English.

Bayanat Labs builds Arabic LLM training data for supervised fine-tuning (SFT): instruction–response pairs, multi-turn dialect conversations, domain-expert Q&A and rewrites. Vetted native speakers and licensed domain professionals write it in Arabic from the start. It is for model builders, banks, telcos, government teams and startups whose model has to answer in the dialect, register and cultural frame their users bring. We work on data only, so we never compete with your model. We specify every set by variety and domain, calibrate it against a gold set, deduplicate it, and deliver it with a datasheet.

25+Arabic varieties
5stages, vetting to delivery
~2 wkstypical pilot

What Arabic LLM training data we deliver

SFT instruction pairs

Single-turn prompts with gold responses for explain, summarize, extract, classify, draft and reasoning tasks. Length, format and register follow your style guide.

Multi-turn dialect conversations

Users write in Gulf, Egyptian, Levantine, Iraqi or Maghrebi Arabic, switch to English or Arabizi, and ask follow-ups that depend on earlier turns.

Domain-expert Q&A

Medical, legal, financial, religious and technical answers, written or reviewed by physicians, lawyers, bankers, engineers and editors.

Rewrites and response repair

We take your model's outputs that have the wrong dialect, translationese, factual slips or off-policy tone, and rewrite them into the response it should learn.

Safety and refusal examples

Adversarial and sensitive Arabic prompts, each paired with the safe completion your policy requires. They come from our Arabic red-teaming work.

Preference data hand-off

After SFT, the same trained writers rank outputs for DPO or reward modeling under one consistent rubric. See Arabic RLHF.

Native vs translated vs synthetic Arabic instruction data

Most Arabic instruction data in circulation began in English. Translation cannot add knowledge the source never had, and it introduces defects of its own. Chen et al. (EMNLP 2024) compared native and translated instruction data. Native and generative benchmarks showed a clear gap, especially for stronger models, but other test sets missed it. The CIDAR team had human reviewers localize 9,109 AlpaGasus pairs machine-translated with ChatGPT. About 64.5% of the pairs needed modification, and the most frequent names shifted from John and Mary to Muhammad and Sarah. TII's Falcon-Arabic announcement (May 2025) says its continued pretraining used 100% native Arabic datasets and avoided machine-translated content.

An illustrative case shows the difference. The English seed "Write a short email to your landlord asking him to fix the heating" translates into correct MSA: اكتب بريدًا إلكترونيًا قصيرًا إلى مالك العقار تطلب فيه إصلاح التدفئة. A Riyadh user types something else: ابي رسالة واتساب للمؤجر اقول له المكيف خربان من امس ("I want a WhatsApp message to the landlord saying the AC has been broken since yesterday"). The channel and the appliance are different, the user writes the Gulf colloquial ابي for "I want", and there is no hamza. A model tuned only on the translated version learns to serve a user who does not exist.

Native-authoredMachine-translatedSynthetic (LLM-generated)
Dialect and registerMatches the target varietyCollapses to MSADrifts to the generator's dominant variety
Cultural groundingLocal names, institutions, normsSource-culture names, holidays, unitsGeneric or stereotyped
Typical defectsWriter inconsistency, caught by gold-set QATranslationese, literal idiomsRepetition, low diversity, confident errors
Cost per exampleHighestLowLowest
Main riskSlower to scaleServes a user who does not existCollapse in the tails, where dialect sits
Best useGold seed, dialect and domain cellsFormat scaffolding, then native localizationExpanding native seeds, behind a native filter

How to design an Arabic training-data mix by dialect and domain

An Arabic SFT set is a matrix. Before writing starts, we fix five things with you:

  1. Varieties and shares. Where you can share samples of real user traffic, we identify their dialects and set quotas per variety, including code-switched and Arabizi turns. See Arabic dialect data.
  2. Response register. Users write in dialect. Your product decides whether the assistant replies in MSA, replies in dialect, or mirrors the user, and the style guide makes that choice a rule.
  3. Domains and owners. Each domain cell names who writes or signs off, for example a banker for fee disputes or a physician for symptoms.
  4. Task types and turn depth. A 2025 review of 366 Arabic post-training datasets found that translation and Q&A make up 42.3% and 38.3% of them. It counted six dialogue datasets and none for function calling, code generation or persona ownership. Commissioned data pays off in those gaps.
  5. Token budget. English-first tokenizers split Arabic into more tokens per word. Measure fertility on your tokenizer before you size the set. See the Arabic tokenization tax.

How a project runs

  1. Scope and vet. We agree the mix with you. Contributors must pass dialect and domain screening for the exact varieties in scope before they touch your project.
  2. Guidelines and gold set. We co-write the style guide, covering register, length, formatting and refusal policy. We also author gold reference examples for writers to calibrate against.
  3. Pilot. A fixed-scope slice of the matrix, checked against the gold set, with a quality report, usually within two weeks.
  4. Produce and QA. Native reviewers score output against the rubric. Disagreements go to adjudication, and failed items return for rework with the reasons attached.
  5. Deliver and scale. You receive JSONL in your chat template with per-item metadata, a datasheet and an audit trail. We then scale as a managed, embedded or enterprise engagement.

Quality, deduplication and documentation

We don't quote a generic accuracy number, because SFT quality only means something against your rubric. For each variety and domain, we run these checks:

  • Gold-set calibration of writers and reviewers before and during production. Drift triggers recalibration.
  • Agreement tracking on double-reviewed items. Low inter-annotator agreement usually means the guideline is ambiguous, so we fix the guideline.
  • Arabic-aware deduplication. We normalize alef variants (أ إ آ to ا), ى/ي, ة/ه, tatweel and diacritics before near-duplicate matching. Without that step, one prompt can survive in two spellings. Lee et al. (2021) found that models trained on deduplicated data emit memorized text ten times less often.
  • Contamination checks against the public benchmarks and internal evals you will report on. The same study found train–test overlap affecting over 4% of the validation set of standard datasets.
  • A datasheet modeled on Datasheets for Datasets. It records composition by variety, domain and task, how items were written and reviewed, known gaps and intended uses.

Rights, consent and data residency

Downloaded data is weakest on provenance. The Data Provenance Initiative found license omission of 70%+ and license error rates of 50%+ on widely used dataset hosting sites. Commissioned data avoids that problem. It is written for your project, the usage terms are agreed before writing starts, and the audit trail records who wrote and reviewed each item. Any documents or logs used for grounding stay in-region by default. On-prem and private-cloud options are available, with least-privilege access and audit logs, and contributors work under NDA. For the regulatory side, see our PDPL and data residency checklist.

Arabic fine-tuning data by industry

  • Finance and banking: fee questions, complaints and compliant refusals in Gulf dialects, reviewed by bankers.
  • Healthcare: patient questions with physician review and clear escalation behavior.
  • Government: citizen queries asked in dialect and answered in a formal register.
  • Legal: clause explanations written by lawyers, with the jurisdiction stated.
  • Telecom: multi-turn support threads with code-switching and Arabizi.
  • Media and retail: rewriting, summaries and product copy in a consistent house style.

When to generate Arabic training data and when to write it by hand

Generation is cheap and works well for format variety, paraphrasing native seeds and red-team breadth. It is weakest on dialect, which is where Arabic SFT data earns its keep. In LabelBench, our reproducible audit, a local language model (Qwen 2.5 3B) reached 22.4% exact-match accuracy on five dialect regions on the fixed held-out manifest, where chance is 20%, and labeled 75% of Gulf examples as Egyptian. Arabic-specialized Fanar 1 9B reached 47.8%. A model that cannot tell Gulf from Egyptian cannot write or filter "Gulf" data. We therefore have native writers author the seed and every dialect and domain cell, and use generation only to expand format. Native reviewers accept or reject each generated item. Model collapse speaks MSA sets out the full argument.

In-house team, crowd platform or specialist vendor?

In-house teamCrowd platformSpecialist Arabic vendor
Dialect-matched writersWhoever you can hireSelf-reported, hard to verifyScreened per variety
Domain expertsStrong in your core domainRareRecruited per domain cell
QA and adjudicationCosts your reviewers' timeOften majority voteGold-set scoring, adjudication, rework
DocumentationAs good as your processOften thinDatasheet and audit trail
ResidencyFull controlWorkers anywhereSet by contract; ask where writers sit
Best whenSmall, sensitive, core-IP setsSimple, low-risk, high-volume tasksDialect, domain and multi-turn data at scale

Start with a pilot

Tell us the model, varieties, domains and task types. We return a fixed-scope pilot set with gold-standard QA and a quality report, usually within two weeks. You can then fine-tune on it, test on your own held-out Arabic prompts, and decide whether to scale based on the results.

Frequently asked questions

How much Arabic SFT data do you need to fine-tune an LLM?

Usually less volume than teams budget for, but more coverage cells than they expect. LIMA aligned a 65B model with 1,000 carefully curated examples, which shows that curation matters more than count. Arabic adds cells, because every variety, domain and task type needs its own coverage. Fine-tune on a pilot slice and measure on held-out native prompts before you scale.

Can we translate English instruction datasets into Arabic instead?

Translation works as scaffolding but rarely as the finished dataset. It carries over the source's topics, names and assumptions, and it cannot produce dialect. A workable compromise is to translate and natively localize the generic MSA tasks, author dialect and domain cells from scratch, and evaluate on prompts written natively in Arabic. See how to benchmark an Arabic LLM.

What is the difference between SFT data and RLHF preference data?

SFT data pairs a prompt with the ideal response and teaches format, register and behavior by example. Preference data records which of several responses is better and why, and it trains a reward model or feeds DPO. SFT usually comes first. We use one style guide for both stages so the model learns one definition of a good answer. See Arabic RLHF.

Which Arabic dialects can you write training data in?

We cover 25+ Arabic varieties with vetted native writers, including Gulf (Saudi, Emirati, Kuwaiti), Egyptian, Levantine, Iraqi and Maghrebi Arabic, plus MSA. We specify coverage at the level your users need, because a single "Gulf" bucket is often too coarse, and we report quality per variety. Our article on dialect coverage explains why this matters.

What format is Arabic training data delivered in?

Usually JSONL in your chat template, with system, user and assistant turns, or any schema your training stack expects. Each item carries its variety, domain, task type, turn count, author role and review status, so you can filter, rebalance or run ablations by slice. A datasheet and an audit trail come with each delivery.

Scope an Arabic SFT pilot

Tell us the dialects, domains and task types your model needs. We come back with a fixed-scope pilot set, gold-standard QA and a quality report, usually within two weeks.

Scope a pilot