Service · Annotation

Arabic data annotation services: labeling for text, speech and documents

Span, intent, dialect, speech and document labels produced by vetted native speakers and domain professionals, measured against gold standards, and kept in-region by default.

Bayanat Labs provides Arabic data annotation services for teams that train, fine-tune or evaluate AI on real Arabic. We deliver entity and span labels, intent and sentiment, classification and moderation, dialect identification, speech transcription, and document and OCR annotation. The work is done by vetted native speakers and licensed domain professionals across 25+ Arabic varieties, calibrated against gold standards, adjudicated, and hosted in-region by default. Work starts with a fixed-scope pilot on your own data, with a quality report usually within two weeks.

25+Arabic varieties and dialects
4modalities: text, audio, image, video
~2 wksusual time to a pilot quality report

What our Arabic data annotation services cover

NER & span labeling

Flat or nested entities, spans and relations, to your schema. بكرة رايح لجدة عشان مقابلة في جامعة الملك عبدالعزيز ("Tomorrow I'm off to Jeddah for an interview at King Abdulaziz University") hides three guideline decisions, shown below.

Intent & sentiment

Intents, sentiment, emotion and sarcasm in dialectal chat and reviews. In Egyptian الخدمة حلوة أوي، بقالي ساعة مستني ("The service is really great, I've been waiting an hour"), the words are positive and the message is a complaint.

Classification & moderation

Multi-label policy and toxicity taxonomies, topics and document classes. تستاهل can mean "you deserve it" or "serves you right". Only the thread decides, so annotators see the thread.

Dialect ID & normalization

وش تبي؟, شو بدك؟ and عايز إيه؟ all mean "what do you want?" (Gulf, Levantine, Egyptian). Many texts fit several varieties: in Keleg and Magdy's error audit, about 66% of a dialect-ID model's validated errors were not true errors. Tags may need to be multi-label. See Arabic dialect data.

Transcription & diarization

Transcripts, timestamps, speaker turns, accent and emotion tags. In خلني أرسل لك الـ invoice على الإيميل ("let me email you the invoice"), the convention decides whether "invoice" stays in Latin script. See Arabic speech data and transcription.

Documents, OCR, image & video

Arabic-script OCR, bounding boxes, segmentation, document fields, and video action and scene tags. A Hijri date such as ١٤٤٦/٠٣/١٢هـ can be captured both as written and as a normalized value. See Arabic OCR and document annotation.

Preference ranking, rewriting and reward-model data sit under Arabic RLHF and human feedback data. To test a finished model rather than label data, see Arabic LLM evaluation services.

Worked example: one Saudi sentence, three span decisions

DecisionOption AOption BWhy it matters
CliticsLOC = لجدةLOC = جدة, prefix لـ outsideOffsets change and must match your tokenizer. The Wojood corpus treats prefixes and suffixes as part of entity names; a pipeline that segments clitics before tagging needs the opposite rule.
NestingORG = جامعة الملك عبدالعزيزAlso PER = الملك عبدالعزيز inside itA flat schema silently drops the person. In Wojood, 22.5% of about 75,000 entities are nested.
Spellingعبدالعزيز (one token)عبد العزيز (two tokens)Normalize or preserve, but decide before measuring agreement, or consistent annotators will look inconsistent.

Arabic dialect and domain coverage

Coverage spans 25+ Arabic varieties and dialects, from MSA to regional and national varieties across the Gulf, Levant, Egypt, Iraq and the Maghreb. Scope a project at the variety level: a Saudi project should name Najdi or Hijazi, not just "Gulf". In 8 of the 12 non-dialect-ID datasets studied by Keleg, Magdy and Goldwater (ACL 2024), covering sentiment, sarcasm, hate speech and stance, full annotator agreement fell as samples became more dialectal. They recommend routing highly dialectal items to native speakers of that dialect. Bergman and Diab (2022) warn that many speakers have only passive knowledge of other varieties, and that many off-the-shelf proficiency tests are specific to MSA. In their case study, country experts judged 74% of the Arabic samples in a Saudi queue, sorted mainly by IP, to be non-Saudi. Routing by geography is not routing by dialect; our guide to dialect coverage explains the wider gap.

Domain matters as much as dialect. Medical, legal, financial and technical judgments go to licensed physicians, lawyers, bankers and engineers; editorial work goes to professional editors.

How an Arabic annotation project runs

Scoping comes first: task, taxonomy, varieties, domain, volume, output schema and residency constraints, with the real dialect mix checked on a sample of your data. Every project then follows the same method: Vet → Train → Produce → QA → Deliver.

  1. Vet. Contributors pass dialect and domain screening for the varieties and subject matter in your data before touching it.
  2. Train. A guideline built around Arabic edge cases (clitics, nesting, sarcasm, Arabizi, code-switching), and a gold set drawn from your data and adjudicated by linguists. Annotators calibrate on gold tasks before production (see our playbook for annotating Arabic text).
  3. Produce. First the pilot: a fixed slice labeled independently, agreement measured, errors analyzed, and the guideline revised wherever disagreement is systematic. Then production at the agreed volume.
  4. QA. Hidden gold items in production queues, overlap sampling, adjudication of disagreements, and rework loops for items that fail review.
  5. Deliver. Structured data in your schema, with quality reporting and an audit trail on every task.

How we measure Arabic annotation quality

  • Gold standards from your data. Adjudicated reference items, not generic tests, used for calibration and as hidden production checks.
  • Chance-corrected agreement. Cohen's kappa for annotator pairs; Krippendorff's alpha for more annotators, missing labels or ordered categories (see Artstein and Poesio's survey); span-level F1 for entity boundaries.
  • Reported per variety and per label. A blended score can hide one failing variety. Targets are set per task in the pilot: a reasonable kappa for sarcasm is not a reasonable kappa for date extraction.
  • Adjudication and rework. Disagreements go to a native linguist. Repeated disagreement changes the guideline, not the annotator, and failed items are reworked and re-checked.

Data security and in-region residency

Data is hosted in-region by default, with on-prem and private-cloud options when it cannot leave your environment. Access is least-privilege, actions are logged, and contributors work under NDA. Annotation is often the first time raw production data reaches a third party, so ask any vendor where the people who open each record sit. Our PDPL and data residency checklist lists the other questions. Saudi buyers can also see data annotation in Saudi Arabia.

Can an LLM pre-label your Arabic data? What LabelBench shows

LabelBench, our reproducible audit (v0.1, July 2026), tests this on the fixed held-out manifest. A local language model, Qwen 2.5 3B, scores 78% on binary Arabic toxicity, but exact-match accuracy falls to 25.3% on the real seven-label taxonomy. On five dialect regions it reaches 22.4% (chance: 20%), and it calls 75% of Gulf examples Egyptian. Arabic-specialized Fanar 1 9B reaches 47.8% on the same test, while a small character model trained on 100 examples per region reaches 69.8%: native, task-specific data still matters more. Under the confidence method tested in v0.1, no generative run could safely automate any share of the queue at a 95% quality floor.

  • Pre-label simple, well-defined labels, and only once the model clears your quality floor on your own gold set.
  • Keep humans on multi-label taxonomies, dialect ID, sarcasm, domain judgments and safety labels.
  • Decide with Safe Automation Coverage: the share of the queue that can be accepted while the 95% Wilson lower confidence bound on accuracy stays above your target.

Arabic annotation use cases by industry

By industry: complaint intents in dialectal banking chat and KYC documents; clinical entities reviewed by physicians; citizen-service intents and document digitization in government; clause and party spans in contracts; call-center transcription and diarization in telecom; review sentiment and moderation in media and retail.

In-house team vs crowd platform vs specialist vendor

CriterionIn-house teamCrowd platformSpecialist Arabic vendor
Dialect matchLimited to the people you can hireSelf-reported, hard to verify per varietyScreened and routed per variety
Domain expertsPossible, costly to keep busyRareMatched to the task
QA methodYou design and staff itUsually platform checks or majority voteGold sets, agreement, adjudication
ResidencyFull controlOften a global workforce; ask where annotators sitVaries; ask for in-region and on-prem options
Ramp-upSlowest: hire and calibrate per varietyFastestA pilot, then scale
Best forStable, long-running core labelsHigh-volume simple labels in well-covered varietiesMulti-dialect, regulated or ambiguous tasks, and gold sets

Comparing vendors? Our Arabic annotation vendor checklist and scorecard turns these criteria into RFP questions.

Start with a fixed-scope Arabic annotation pilot

A pilot covers one task, an agreed sample of your data and named varieties. It includes gold-standard QA and ends with a quality report, usually within two weeks: agreement per variety and label, error analysis, guideline revisions, and how much of the queue could safely be pre-labeled. To scope one, send a representative sample (masked where the task allows), your taxonomy, target markets, residency constraints and output format through our contact page. Production dates are set after the pilot, once throughput on your data is measured, and work scales as a managed, embedded or enterprise engagement.

Looking for annotation work? This page is for teams commissioning data. Native Arabic speakers can apply free for remote, paid-per-task work through the Bayanat Talent Network.

Frequently asked questions

How much do Arabic data annotation services cost?

Cost follows effort per item, not the language. The drivers are taxonomy size, how many varieties the data mixes, whether a licensed expert must judge each item, how much QA overlap is needed, and modality: an hour of dialectal call audio takes far longer than a social post. We quote after the pilot measures throughput. See our breakdown of Arabic annotation cost drivers.

Is there a difference between data annotation, labeling and labelling?

Labeling (American) and labelling (British) are the same word, and most buyers use either interchangeably with annotation. Where teams draw a line, labeling assigns a class to a whole item, while annotation also marks structure inside it: entity spans, relations, timestamps, speaker turns, bounding boxes. Our Arabic data labeling guide covers each task type.

How do you handle Arabizi and code-switched Arabic?

As part of the guideline, not as noise. Before production we agree whether Arabizi such as 3ayez (Egyptian for 'I want') is labeled as written or also transliterated into Arabic script, whether English or French inserts get a language tag, and how entity spans cross a script boundary. The gold set then tests those rules like any other edge case.

How many annotators label each item?

It depends on how subjective the task is. Objective tasks such as date or amount extraction can run with one annotator, hidden gold checks and an overlap sample for measuring agreement. Subjective tasks such as sarcasm, toxicity or dialect ID need two or more independent native annotators, with disagreements adjudicated. The pilot shows which design your error rate justifies.

Scope a pilot on your own data

Tell us the dialects, domain, modality and volume. We come back with a fixed-scope pilot, gold-standard QA and a quality report — usually within two weeks.

Talk to us