Checklist · Vendor selection

How to choose an Arabic data annotation vendor: RFP checklist and scorecard

Generic outsourcing checklists ask about price and turnaround. For Arabic, the questions that decide quality are different: who is each rater, where do they sit, and what evidence shows the labels hold up in every dialect you ship.

Bayanat Labs Research· · 11 min read

To choose an Arabic data annotation vendor, test what generic outsourcing checklists skip. Can the vendor name and verify each rater's variety? Where do the people who open your records sit? Can it evidence quality per dialect, prove its experts are licensed, and pilot on your data against criteria you set? Send the 20 RFP questions below, apply the knockout gates, score the answers, and let the pilot decide.

This page is vendor-neutral. Bayanat Labs is itself an Arabic annotation vendor, and every question applies to us too. If you are still deciding whether to outsource at all, start with our Arabic data labeling guide.

What separates an Arabic annotation specialist from a generalist?

The difference shows when you ask for evidence instead of assurances.

AreaA generic answer (illustrative)Evidence to request
Dialect provenance"All our annotators are native Arabic speakers."Per-rater variety profiles and dialect screening results
Quality evidence"99% accuracy."Gold-set design, chance-corrected agreement by dialect, an adjudication log
Rater location"Your data is stored securely in the cloud."The countries where every person who opens a record sits; a full sub-processor list
Domain experts"Subject-matter experts are available."License type, issuing body and a way to verify each license
Automation"An AI-assisted workflow."Which labels are pre-labeled, and how safety was measured on your taxonomy
Pilot terms"A sample on our demo data."A fixed-scope pilot on your data, with acceptance criteria written first

Dialect provenance: can the vendor name each rater's variety?

"Native Arabic speaker" is not a qualification for dialectal data. Keleg, Magdy and Goldwater (ACL 2024) examined 15 public Arabic datasets with individual annotations. In 11, the more dialectal a sample, the less its annotators agreed; the authors recommend routing highly dialectal items to native speakers of that dialect. Bergman and Diab (2022) add that many speakers have only passive knowledge of other varieties, and that many off-the-shelf Arabic proficiency tests measure MSA. Passing one proves little about a Gulf queue.

One word shows why. مرة (marra) means "once" in MSA and "very" in Saudi speech: مرة حلو, "really nice". In Levantine, the same spelling is a common way to write "woman" or "wife" (mara). Read the wrong sense and sentiment, intent and entity labels all go wrong. Even inside Saudi Arabia, a Najdi speaker asks وش (wesh, "what") where a Hijazi speaker asks إيش (ēsh). A Saudi project needs raters named at that level, not just "Gulf".

So ask for per-rater profiles, a dialect-based screening test, routing by variety and multi-label dialect tags (questions 1–4 below). When native speakers reviewed a dialect-ID model's errors, Keleg and Magdy (2023) found about 66% of validated errors were not true errors, and argue that dialect identification should be multi-label. Our dialect coverage guide helps you choose which varieties to require.

Quality evidence: gold sets, agreement, adjudication and audit trails

An accuracy figure means nothing until you know what it was measured against. Ask for four artifacts, redacted from a past project or produced in your pilot:

  1. Gold set design. Who wrote it, how it is stratified by dialect and label, and what share of production items are hidden gold.
  2. Chance-corrected agreement. Cohen's kappa for pairs, Krippendorff's alpha for several raters, span F1 for entities, reported by label and by dialect. Raw percentage agreement flatters imbalanced tasks.
  3. An adjudication log. Who broke each tie, the rationale, and which decisions changed the guideline.
  4. An audit trail. Rater, reviewer, timestamps, edits and guideline version for every item.

Be wary of generic thresholds. Artstein and Poesio's survey notes that the field's customary 0.8 cutoff was borrowed from content analysis, and doubts that one cutoff suits every annotation task. Set your floor per task in the pilot, then write it into the contract. Our Arabic text annotation playbook covers the guideline decisions that most often move agreement.

Data residency: where do the humans who open each record sit?

Server location is the easy question. The revealing one is where the people who open each record sit, and under whose contract. For Saudi personal data, that is your obligation too. PDPL Article 8 requires controllers to select only processors that provide the necessary guarantees, and to monitor them. Article 17 of the Implementing Regulation says the processor agreement must identify subcontractors, state whether the processor is subject to other countries' regulations, and commit to breach notification. It is also sensible to require your prior approval of any new sub-processor, with time to object.

Each clause becomes an RFP question (11–15 below). Crowd routing is where residency most easily breaks, so ask about platforms and marketplaces by name. If you sell high-risk AI systems into the EU, Article 10(2) of the EU AI Act also requires data governance covering data origin and preparation operations "such as annotation, labelling", so your annotation records must be documentable. See our PDPL and data residency checklist for transfer rules. This is general information, not legal advice.

Domain experts for medical, legal and financial Arabic data

The question is not whether a vendor has experts, but whether you can verify them. Ask for each expert's license type and issuing body. In Saudi Arabia, for example, the Saudi Commission for Health Specialties offers an online check of a practitioner's registration. Then ask who adjudicates domain disputes: a licensed professional or a linguist.

Experts also need the right variety and jurisdiction. A Gulf patient says يعورني (yʿawwirni, "it hurts me"); an Egyptian patient says بيوجعني (biyewgaʿni). Saudi statutes are usually titled نظام (niẓām), as in نظام حماية البيانات الشخصية, the PDPL, while Egypt's data protection law is a قانون (qānūn). A reviewer who knows law but not the country will miss references like these.

Machine pre-labeling: ask where automation is actually safe

Model pre-labels with human correction can work, but only where they are measured on your task. In v0.1 of LabelBench, our reproducible audit, a small zero-shot model (Qwen 2.5 3B) scored 78% on binary Arabic toxicity, then 25.3% exact-match on the real seven-label taxonomy and 22.4% on five-region dialect identification. Under the confidence method tested, no generative run could safely automate any share of the queue at a 95% quality floor. That reflects v0.1's routing protocol, not proof that automation never works, but a model can look fine on a simple task and fail on yours. Require the pilot to report human-only and assisted items separately.

Pilot terms that de-risk the decision

An RFP shows what a vendor says; a pilot shows what it delivers. Agree these terms first:

  • Your data, stratified by dialect, channel and difficulty, including your hardest edge cases. A few hundred items is often enough (illustrative; scale with your label count).
  • Hidden gold the vendor cannot tune to, and written floors for agreement and gold accuracy per dialect.
  • The production team: named raters who stay on at scale.
  • Deliverables beyond labels: a per-dialect quality report, error taxonomy, adjudication log, measured throughput and guideline-question log.
  • Exit terms: deletion on completion, ownership of the guideline you co-wrote, and scale-up prices agreed in advance.

To turn pilot results into comparable quotes, see what drives Arabic data annotation cost.

Arabic annotation RFP template: 20 questions

  1. Dialect. What is the native variety of each proposed rater (country, and region or city where relevant)? How did you verify it beyond self-report?
  2. Dialect. Which screening test do raters pass before touching client data? Does it test dialect comprehension, not only MSA? Share a sample item.
  3. Dialect. How do you route items by variety, and what happens to items whose variety is unclear or mixed?
  4. Dialect. Can dialect tags be multi-label for items valid in several varieties?
  5. Quality. Who builds the gold set, how is it stratified, and what share of production items are hidden gold?
  6. Quality. Which agreement metrics do you report: per label, per dialect, per rater?
  7. Quality. Who breaks ties in adjudication, how are rationales logged, and how do decisions update the guideline?
  8. Quality. What does your audit trail record per item? Share a redacted export.
  9. Quality. How do you detect raters drifting below the gold threshold, and is their earlier work re-checked?
  10. Guidelines. How will you handle hamza and alef forms, taa marbuta, diacritics, clitics in spans, Arabizi and code-switching? Show a sample guideline page.
  11. Residency. In which countries do the people who will open our records sit, including contractors and crowd workers?
  12. Residency. List every sub-processor, including labeling platforms and crowd marketplaces. Will you notify us before changes and let us object?
  13. Residency. Where are data stored and processed? Can you work in-region, inside our environment or on-premises?
  14. Residency. Are you subject to other countries' laws that could compel disclosure of our data?
  15. Security. What access controls apply (least privilege, no local downloads, NDAs)? What are your breach-notification commitment and deletion plan?
  16. Domain. Which licensed professionals review medical, legal or financial items, who issued each license, and how can we verify it?
  17. Automation. Which labels are machine pre-labeled, how was their safety measured on our taxonomy and dialect mix, and what share gets full human review?
  18. Pilot. Propose a fixed-scope pilot on our data: size, dialect strata, acceptance criteria, deliverables and timeline.
  19. Scale. Will the pilot team be the production team? How will you add raters for rare varieties without lowering the screening bar?
  20. Commercials. What are your pricing units, which rework is at your cost, and how does price change if the taxonomy or dialect mix changes?

Arabic data annotation vendor scorecard

Apply three knockout gates first: drop any vendor that will not name the countries its annotators work from, list its sub-processors, or pilot on your data. Score the rest from 1 to 5. The weights are a starting point; adjust them to your risk. Weighted score = sum of (score ÷ 5 × weight), out of 100.

CriterionWeightScores 1Scores 5
Dialect provenance20"Native speakers"; MSA-only testVerified per-rater varieties, dialect screening, routing by variety
Quality evidence20One accuracy figure, no methodStratified hidden gold, agreement by dialect, adjudication log, audit export
Pilot results on your data20Missed the floor, or no written criteriaMet every dialect floor with the production team
Residency and security15Vague rater locations; crowd routingNamed rater locations, in-region hosting, sub-processor consent, deletion plan
Domain expertise10Unverifiable "experts"Verifiable licenses; experts adjudicate domain disputes
Commercials and scale path10Per-label price only; rework billedClear units, rework terms, agreed prices for change and scale-up
Automation transparency5Undisclosed pre-labelingMeasured on your taxonomy; assisted items reported separately

Have two people score each proposal independently before comparing notes: double annotation, applied to procurement.

Red flags in an Arabic annotation proposal

  • Rater qualifications stop at "native Arabic speakers", or at one "Gulf" label for a Saudi project.
  • Accuracy claims with no gold set, no chance-corrected metric and no dialect breakdown.
  • "Our workforce is global" when you ask where raters sit.
  • Sample data requested before an NDA and data processing agreement are signed.
  • A pilot on the vendor's own data, or run by a team that will not do the production work.
  • No written commitment to delete raw data and intermediate files.

How Bayanat Labs answers these questions

Bayanat Labs is a Riyadh-based Arabic data engine covering 25+ Arabic varieties. Our contributors are vetted native speakers, linguists and licensed domain professionals who pass dialect and domain screening before touching client data. Work is calibrated against gold standards, with adjudication, rework loops and an audit trail on every task. Hosting is in-region by default, with on-prem and private-cloud options, least-privilege access and contributors under NDA. Engagements start with a fixed-scope pilot and a quality report, usually within two weeks. See our Arabic data annotation services, or data annotation in Saudi Arabia if your data must stay in-Kingdom.

Key takeaways
  • Ask for evidence, not adjectives. Per-rater dialect profiles, agreement by dialect, adjudication logs and audit exports separate specialists from generalists.
  • Residency is about people. Ask where everyone who opens a record sits, and name every sub-processor.
  • Verify experts and automation. Licenses should be checkable, and pre-labels measured on your taxonomy.
  • Let a written pilot decide. Knockout gates, then a weighted score, then a pilot on your data against criteria set in advance.

Frequently asked questions

How many Arabic annotation vendors should we pilot?

Send the RFP widely, then pilot only the two or three vendors that clear the knockout gates and score highest on paper. Each pilot costs your team review time. Give every finalist the same items, the same hidden gold and the same deadline, so results are directly comparable.

What inter-annotator agreement should an Arabic vendor achieve?

There is no universal number. Binary moderation labels should agree far more often than fine-grained sentiment or entity spans in dialectal text. Use a chance-corrected metric such as Krippendorff's alpha, and set the floor for each dialect separately: a strong overall figure can hide one variety that falls well short.

Should an Arabic annotation pilot be paid or free?

A paid, fixed-scope pilot is usually the better test. Paying lets you require what a free sample may not include: your own data, the team that will do production work, hidden gold items and a written quality report by dialect. Keep the scope small and the acceptance criteria fixed in advance.

Can we send Saudi personal data to a vendor outside the Kingdom?

Saudi Arabia's PDPL regulates cross-border transfers rather than banning them, but each transfer route adds conditions and safeguards, and you stay responsible for the vendor's processing. In-Kingdom annotation removes most of those steps. Our PDPL data residency checklist has the details. This is general information, not legal advice.

What should an Arabic annotation RFP include besides questions?

Attach a redacted data sample, your draft guideline and taxonomy, and the dialect and register mix you expect in production. State volumes, the delivery format and the pilot acceptance criteria. Publish your scoring weights too: vendors answer more precisely when they know what you will measure.

Put these 20 questions to us

Bayanat Labs runs a fixed-scope pilot on your data, with gold-standard QA and a benchmark and quality report, usually within two weeks.

Scope a pilot