Comparison · Vendors

Arabic Data Annotation Companies Compared (2026): A Buyer's Guide

Most lists of annotation vendors ignore Arabic, and most Arabic pages are written by a single vendor about itself. This one lays each company's own published claims side by side, with the same columns, a source link for every fact and a date on every row.

Bayanat Labs Research· · 13 min read

There is no neutral list of the best Arabic data annotation companies. Search results for the query are mostly vendor service pages, vendor-written "top N" lists and general outsourcing directories that rarely compare dialect coverage. This page compares 12 Arabic data annotation companies and dataset sources on the same six criteria, using only what each company states on its own public pages as of September 2026. Every fact links to its source.

The short answer: regional specialists (Alaraby AI, Annota8, Unidata, WASM) build from the Arab world and several name dialects in detail; global vendors (AI Taggers, Annotera, Appen, Pangeanic, Toloka) run Arabic alongside other languages and often list broader modality menus; and off-the-shelf catalogues (Defined.ai, Shaip, TELUS Digital, plus the free LDC and Masader) suit teams that need data now rather than a custom labeling programme. Which is "best" depends on your dialects, modality and residency rules, not on a league table.

Disclosure. Bayanat Labs publishes this comparison and is itself an Arabic data annotation vendor. We list ourselves in a separate, labelled row using the same criteria and only the claims we make on our own site. We do not rank any company, including ourselves, and no vendor paid to appear.

How this comparison of Arabic data annotation companies works

Most "top annotation companies" lists sort firms by reviews, sponsorship or size. None of those tell you whether a vendor can label Moroccan Darija or Najdi speech. So this comparison uses six Arabic-specific criteria, and records each one only as the vendor itself publishes it:

  1. Dialects stated. Which Arabic varieties the page names. "Arabic" alone counts as unspecified.
  2. Modalities and tasks. Text, speech, image, video, documents, and which task types.
  3. RLHF and LLM evaluation. Whether the page offers preference data, human feedback or evaluation of model outputs.
  4. Stated location and residency. Where the company says it is based, and any statement about where data or annotators sit.
  5. Pilot or sample terms. Whatever the page says about trials, samples or pilots.
  6. Published quality evidence. Processes, metrics or certifications the page claims, reported as claims.

Three rules apply throughout. First, "not stated on page" means the linked page does not say it, not that the company does not offer it. Second, facts are as stated on vendors' own pages as of September 2026; Bayanat Labs did not test, audit or rate any vendor, and vendor claims are reported, not endorsed. Third, companies appear alphabetically within each group. Companies whose Arabic offering could not be confirmed on a live page were left out rather than guessed at.

Arabic annotation vendors compared: the criteria table

All rows are as of September 2026. Follow each vendor link to check the current wording before you rely on it.

Company (source page)TypeDialects statedRLHF / LLM evaluationStated locationPilot or sample termsQuality evidence on page
Regional Arabic specialists
Alaraby AIManaged serviceGulf; Saudi Najdi and Hejazi; Levantine; Egyptian; Iraqi; North AfricanLLM output evaluation listed as a taskNot stated on pageNot stated on pagePage states ISO certification in linguistic services, 20 native annotators, 100+ projects; agreement "tracked" (no figures)
Annota8Annotation platformNajdi transcription shown as an example templatePairwise RLHF, ranking, RAG and agent-trace evaluation interfacesBuilt in Saudi Arabia (per site); regions listed KSA, USA, EgyptDemo bookingNot stated on page
UnidataManaged service (case study)Gulf (UAE, Saudi); North African (Morocco, Algeria)Linguistic evaluation of LLM-generated Arabic text (case study)Dubai address on pageCase timeline includes a 1-week pilotNative proficiency validated through test tasks; dialect-matched assignment
WASMManaged serviceArabic and Egyptian ArabicNot stated on pageCairo, EgyptFree scoping call; pilot tasks listed in its processGuidelines, calibration tasks, multi-annotator agreement checks, QA passes
Global vendors with Arabic programmes
AI TaggersManaged serviceMSA plus Gulf, Levantine, Egyptian, Maghrebi, IraqiArabic LLM RLHF and instruction-tuning datasetsMelbourne, Australia; page says it adheres to Saudi PDPL and UAE requirementsFree sample of 25–50 records in 24–48 hours; 1,000-record pilot in 3–5 business daysAustralian-led QA stated; no metrics
AnnoteraManaged serviceMSA, Gulf, Egyptian, Levantine, MaghrebiNot stated on page (generative AI data listed)Not stated on pageNot stated on pageMulti-level quality validation stated; no metrics
AppenManaged service and datasetsArabic named only in an Arabic-French code-switching exampleSite menu lists SME RLHF, SFT and red teaming (not Arabic-specific)Not stated on this pageNot stated on this pagePage states SOC 2 and ISO 27001 certification
PangeanicDatasets and managed serviceMSA; Gulf; Levantine; Maghrebi; Egyptian and Sudanese; Arabic-English and Arabic-French mixingModel alignment and RLHF in site menu; evaluation test setsNot stated on page (footer: Pangeanic S.L.)Samples depend on each dataset's licensing conditionsDialect verification and sample inspection described; no metrics
TolokaEvaluation-data project (blog, not a service page)Jordanian, Egyptian, Emirati, MoroccanJEEM benchmark for vision-language evaluationNot stated on pageNot stated on pageAnnotators passed a qualification test of writing in the target dialect and MSA
Off-the-shelf datasets
Defined.aiDataset marketplaceJordanian, Egyptian, Yemeni, MSANot stated on pageNot stated on pageNot stated on pageHours and locale per dataset
ShaipSpeech data catalogueGulf Arabic named for some conversationsNot stated on pageNot stated on pageNot stated on pageTotal hours and speaker count stated
TELUS DigitalOff-the-shelf speech datasetSaudi Arabic (ar-SA)Not stated on pageNot stated on pageSample and pricing on requestStudio recording specification stated
Publisher of this page (disclosed, not ranked)
Bayanat LabsManaged service25+ Arabic varietiesPreference ranking, SFT and rewriting, reward modeling, red-teaming, evaluationRiyadh; in-region hosting by default, on-prem and private-cloud optionsFixed-scope pilot with a quality report, usually within two weeksGold-standard calibration, adjudication, audit trail; LabelBench audit published

Regional Arabic data annotation specialists

These four companies are based in, or build from, the Arab world. Their pages tend to name dialects more precisely than global vendors do.

Alaraby AI

Alaraby AI's Arabic data annotation page names Saudi Najdi and Hejazi separately from general Gulf Arabic. It also names Levantine, Egyptian, Iraqi and North African varieties, and says Gulf is the dialect most requested in its project history. Its task list runs from classification, NER, sentiment and intent to speech transcription, audio segmentation and LLM output evaluation. The page states that work is done by 20 native annotators, that the company holds an ISO certification in linguistic services and has completed 100+ projects, and that inter-annotator agreement is tracked. It publishes no agreement figures, and does not say where data is processed.

Annota8

Annota8 is a platform rather than a labeling workforce, which makes it a different kind of choice. Its public roadmap lists workforce management as the next module and describes an end-to-end managed pipeline as a longer-term vision. The site says it is Arabic-first and built in Saudi Arabia, lists KSA, USA and Egypt as regions, and offers 200 annotation interfaces across nine data modalities, including pairwise RLHF comparison, RAG relevance rating and agent-trace evaluation. Najdi dialect transcription appears as an example use case. Residency and pricing are not stated on the page.

Unidata

Unidata's Arabic evidence is a case study for an unnamed telecom client: verbatim transcription, audio evaluation and linguistic evaluation of LLM-generated Arabic text, to validate internal AI tools. It covered Gulf (UAE and Saudi) and North African (Moroccan and Algerian) speech, including English loanwords and French insertions. The page says native proficiency was validated "through test tasks, not profiles" and annotators were matched to tasks strictly by dialect. The timeline shows a 1-week pilot inside a seven-week project. The page lists a Dubai address.

WASM

WASM, based in Cairo, focuses on Arabic and Egyptian Arabic alongside Egyptian street computer vision and medical data with doctor review. Its NLP tasks include sentiment, intent, NER, toxicity, search relevance and summarization. It describes its quality process as guidelines, calibration tasks, multi-annotator agreement checks and QA passes, and offers a free scoping call, with guidelines and pilot tasks as the next step in its process. It names no other Arabic dialect, which suits an Egyptian-market project and leaves Gulf or Maghrebi coverage as a question to ask.

Global vendors with Arabic programmes

Global vendors often list broader modality menus. Where Arabic is one language among many, check how deep each named variety goes.

  • AI Taggers (Melbourne) publishes the most concrete trial terms in this comparison: a free 25–50 record sample in 24–48 hours, then 1,000-record production pilots in 3–5 business days. It lists MSA plus five dialect families, naming countries for Gulf, Levantine and Maghrebi, and describes Arabizi handling, and Arabic LLM RLHF and instruction-tuning datasets. The page says it adheres to Saudi PDPL and UAE data protection requirements, but does not state where annotators sit.
  • Annotera covers MSA, Gulf, Egyptian, Levantine and Maghrebi across text, image, audio and video, with sentiment, moderation and generative AI data. Location, pilot terms and quality metrics are not stated on the page beyond a description of multi-level quality validation.
  • Appen's multilingual page names Arabic only as an Arabic-French code-switching example, in a programme it says spans 500+ locales and 80+ languages. The page states that Appen is SOC 2 and ISO 27001 certified. A November 2025 Appen blog post says it uses inter-rater reliability measures such as Krippendorff's alpha to qualify contributors and calibrate reviewers.
  • Pangeanic offers both licensed Arabic datasets and bespoke collection, with annotation through its PECAT platform, evaluation, and model alignment and RLHF. Its dialect list includes Sudanese, and it covers Arabic-English and Arabic-French mixing. It publishes no dataset sizes on this page; sample access depends on each dataset's licence.
  • Toloka is included for evaluation-data evidence, not a service page. With MBZUAI it built JEEM, an image-captioning and visual question-answering benchmark in Jordanian, Egyptian, Emirati and Moroccan Arabic: 2,178 annotated images in 13 categories, according to Toloka's post; the benchmark itself is described in the JEEM paper (March 2025). The post says each annotator passed a qualification test of writing skills in the target dialect and in MSA.

Arabic dataset marketplaces and free catalogues

If you need data this month rather than a labeling programme, an off-the-shelf dataset can be faster. Check three things before buying: whether the dialect matches your users, whether the licence allows commercial training, and whether the recording conditions match production. A studio wake-word set will not teach a model noisy call-centre speech. Our guide to Arabic speech datasets covers these trade-offs.

  • Defined.ai lists simulated call-centre conversations recorded over telephony, in Jordanian (insurance, banking, telco, retail), Egyptian, Yemeni and MSA. The listing shows 20 datasets, with sizes on the first page ranging from 11 to 34 hours each.
  • Shaip's Arabic speech catalogue states a total of 7,894 hours and 2,761+ speakers across call-centre, general conversation, scripted monologue and singing audio. It names Gulf Arabic for its human-to-human telephone conversations. Check each set's hours and dialect separately before you buy.
  • TELUS Digital sells a Saudi Arabic studio speech dataset for wake-word and command recognition: 1,392 prompts from 29 adult participants, 1 hour 4 minutes of 44.1 kHz, 24-bit mono audio. Samples and pricing are available on request. At about an hour of prompted studio speech, it is built for wake-word and command work rather than conversational ASR.
  • The Linguistic Data Consortium (LDC) licenses research corpora. One example, the Arabic-Dialect/English Parallel Text released on June 15, 2012, holds approximately 3.5 million tokens of Egyptian and Levantine Arabic web text with English translations, under the LDC User Agreement for Non-Members.
  • Masader is not a vendor. It is a free public catalogue of Arabic NLP and speech datasets: the repository summary says 500+ datasets and the README more than 600, each described with more than 25 attributes, and a good first stop before paying for anything.

Which type of Arabic annotation vendor fits your project?

The groups above solve different problems. Start from your need, not from a list.

If you need…Look first atAsk before signing
Custom labels in one or two named dialectsRegional specialistsHow many screened raters per variety, and how dialect was verified
Many languages, of which Arabic is oneGlobal vendorsWhich Arabic varieties they staff today, not in principle
Preference data or LLM evaluationVendors that state RLHF or evaluationRater calibration method and agreement on your rubric
Your own team labeling on your infrastructurePlatformsHosting options and who operates the workforce
Speech data this monthDataset cataloguesLicence terms, dialect per file, recording conditions
Saudi personal data kept in-KingdomVendors that state in-region processingWhere every person who opens a record sits

For Saudi personal data, residency is a legal question as well as a commercial one. Our PDPL data residency checklist explains the transfer rules. This is general information, not legal advice.

How to shortlist Arabic data annotation companies in five steps

  1. Write your dialect and modality mix first. Pull a sample of real user text or audio and tag its varieties. Our dialect coverage guide helps you decide which ones matter.
  2. Filter by vendor type. Use the table above to drop groups that cannot deliver your task at all, such as a speech catalogue for NER work.
  3. Turn every "not stated on page" cell into a question. Our Arabic annotation vendor RFP checklist has 20 questions and a weighted scorecard.
  4. Probe dialect claims with real words. "I want" is أبغى (abgha) in much of the Gulf, عايز (ʿāyiz) in Egypt and بدّي (biddi) in the Levant. Put items like these in a short screening set and see whose raters label them correctly.
  5. Pilot two or three finalists on the same data. Give each the same items, hidden gold and acceptance criteria, then compare quotes on the same units. Our breakdown of what drives Arabic data annotation cost shows how to normalise them.

Ask every vendor how much of the work is machine pre-labeled, too. In our LabelBench v0.1 audit, a local language model that reached 78% on binary Arabic toxicity fell to 25.3% exact-match accuracy on the real multi-label taxonomy, and scored 22.4% across five dialect regions (chance 20%), on the fixed held-out manifest. A model can look fine on a simple task and fail on yours, so assisted labels should be measured on your taxonomy and reported separately.

Where Bayanat Labs fits

Bayanat Labs, the publisher of this page, is a Riyadh-based Arabic data company. We do data only, so we never compete with a client's models. We cover 25+ Arabic varieties with vetted native speakers, linguists and licensed domain professionals who pass dialect and domain screening before touching client data. Work is calibrated against gold standards, with adjudication, rework loops and an audit trail on every task. Hosting is in-region by default, with on-prem and private-cloud options. Engagements start with a fixed-scope pilot and a quality report, usually within two weeks. If your data must stay in the Kingdom, see data annotation in Saudi Arabia; otherwise, start with our Arabic data annotation services. Hold us to the same questions as everyone above.

Key takeaways
  • No neutral ranking exists. Compare vendors on dialects, tasks, RLHF and evaluation, location, pilot terms and published quality evidence.
  • Pick the vendor type first. Regional specialists, global vendors, platforms and dataset catalogues solve different problems.
  • Blank cells are RFP questions. Most public pages skip residency, pilot terms and measured quality, so ask for documents.
  • Let a pilot decide. Run the same items, hidden gold and criteria past two or three finalists.

Frequently asked questions

Is this list of Arabic annotation companies sponsored?

No. No company paid to appear. The page is published by Bayanat Labs, which is itself an Arabic annotation vendor, so we list ourselves in a separate, labelled row and do not rank anyone. If a vendor's page changes, or a vendor believes a fact is out of date, email hello@bayanatlabs.com with the live source.

Are there Arabic data annotation companies based in Saudi Arabia or the UAE?

Yes. As of September 2026, Unidata's case study gives an address in Dubai, Annota8 describes its platform as built in Saudi Arabia, and Bayanat Labs is based in Riyadh. A head office is not the same as where the work happens, though. Ask every vendor, wherever it is based, which countries its annotators and reviewers work from.

How many Arabic dialects should an annotation vendor cover?

Only the ones your users speak, but covered deeply. A vendor listing six dialect families helps little if your Saudi product needs Najdi and Hijazi raters and the vendor can offer only generic Gulf. Pick the varieties from your own traffic or call logs first, then ask each vendor for the number of screened raters in each one.

Why do so many cells say "not stated on page"?

Because this comparison records only what each company publishes, and most public pages skip data residency, pilot terms and measured quality. That does not mean the vendor lacks them. It means you must ask. Turn each blank cell into a written RFP question, and prefer vendors that answer with documents rather than adjectives.

Can one vendor handle Arabic speech, text and LLM evaluation?

Several say so on their own pages, but the skills differ. Speech work needs transcription conventions and acoustic judgement, while LLM evaluation needs raters who can rank fluent answers for factual and cultural errors. If one vendor covers both, ask whether the same team does both, and pilot each task type separately.

Put us through the same questions

Bayanat Labs runs a fixed-scope pilot on your data, with gold-standard QA and a benchmark and quality report, usually within two weeks.

Scope a pilot