Service · Saudi Arabia

Data annotation company in Saudi Arabia: Arabic-native, in-region by default

Labeling, collection, alignment and evaluation data for AI teams in the Kingdom: Saudi varieties handled by screened native speakers, domain work by licensed professionals, and data handling you can explain to SDAIA and your sector regulator.

Bayanat Labs is a data annotation company in Saudi Arabia, based in Riyadh. We label, collect, align and evaluate Arabic data for AI across text, audio, image and video, using vetted native speakers, linguists and licensed domain professionals, and we host in-region by default. The service is built for Saudi model builders, government entities, banks, telcos and healthcare organizations that need two things at once: Saudi Arabic labeled by people who speak it, and data handling they can explain to SDAIA and to their sector regulator.

25+Arabic varieties and dialects
4modalities: text, audio, image, video
~2 weekstypical fixed-scope pilot

Why use a data annotation company in Saudi Arabia?

Annotation is the step where raw data is most exposed. People open every call recording, chat log and scanned form in full. Three things make that step hard to send to an offshore crowd:

  • The PDPL follows the data. It covers processing of personal data about people residing in the Kingdom, including processing done by a party outside it (PDPL, Article 2). You must choose processors that give sufficient guarantees (Article 8), and the processor agreement must state whether the processor is subject to other countries' regulations (Implementing Regulation, Article 17).
  • Sector rules add friction offshore. SAMA's Rules on Outsourcing require banks to get a written SAMA no-objection before outsourcing to a provider located overseas, and the request must explain why the work cannot be done in the Kingdom (paragraph 42). SDAIA's data management and personal data protection standards also reach business partners that handle government data.
  • Saudi Arabic is not a single variety. The SDC-5 corpus divides Saudi dialects into five regional groups: Hijazi, Najdi, Southern, Northern and Eastern (Al-Shenaifi et al., 2024). An annotator pool recruited generically for "Gulf Arabic" misses differences that your users produce every day.

What we deliver in-region

Text annotation and labeling

NER and span labeling, intent, sentiment, classification and moderation on dialectal, code-switched and Arabizi text. See Arabic data annotation services.

Speech data and transcription

Transcription, timestamping, speaker diarization, and accent and emotion tagging for call-center and voice-assistant audio. See Arabic speech data.

Dialect data

Dialect identification, targeted collection and dialect-balanced sets for when Najdi or Hijazi traffic is thin in your data. See Arabic dialect data.

Documents, OCR, image and video

Arabic-script OCR ground truth, document understanding, bounding boxes and segmentation for forms, contracts and scanned records, plus action, tracking and event labels for video. See Arabic OCR and document annotation.

LLM training and alignment data

SFT and instruction data, rewrites, preference ranking and reward-model data written and judged by native speakers. See Arabic LLM training data and Arabic RLHF.

Evaluation and red teaming

Evaluations for dialect comprehension, cultural fit, factuality and safety, plus adversarial testing in dialect and Arabizi. See Arabic LLM evaluation and Arabic AI red teaming.

Which Saudi dialects do you cover?

Najdi, Hijazi and the other Saudi regional varieties are part of our 25+ variety coverage. We agree the exact mix at scoping, and contributors pass screening for the specific variety before they touch client data. The differences are more than accent. The same request is worded differently across the Kingdom:

Variety"What do you want?""Now"What it means for labeling
MSAماذا تريد؟ (mādhā turīd?)الآن (al-ʾān)Official documents, forms and news. Mostly OCR and document work.
Najdi (central)وش تبي؟ (wesh tabi?)الحين (al-ḥīn)Chat, voice notes and call audio. Intent lexicons built on MSA miss these words.
Hijazi (western)إيش تبغى؟ (ēsh tibgha?)دحين (daḥīn)The same intents in different words. A model or guideline tuned on Najdi under-labels them.
Saudi–English code-switchingابي الغي الـ subscription حقي ("I want to cancel my subscription")The intent spans two scripts, so guidelines must define how to tag mixed-script spans.

Speech is harder still. SADA is a public corpus of 668 hours from 57 Saudi Broadcasting Authority television shows, primarily Najdi, Hijazi and Khaleeji speech. In a 2025 study, the best system (MMS 1B fine-tuned on SADA with a 4-gram language model) still had a 40.9% word error rate on the SADA clean test set (arXiv:2508.12968). For a way to measure coverage gaps in your own data, see why dialect coverage decides your Arabic model's ceiling.

PDPL and data residency: what to ask any annotation vendor

This section is general information, not legal advice. Regulations and guidance change; confirm current requirements with counsel and with SDAIA's published texts before you make compliance decisions.

No vendor can make your project "PDPL-compliant data annotation" on your behalf, because you remain the controller. What a vendor can do is give answers you can verify. The PDPL does not ban transfers: Article 29 permits them under conditions, and SDAIA's Transfer Regulation (v2.0, August 2024) lists safeguards such as standard contractual clauses and binding common rules, plus a risk assessment in defined cases. Processing in-Kingdom or in-region removes most of those steps.

Ask the vendorWhy it matters
Where are storage, processing and the annotators located?This decides whether the transfer rules apply at all. Server location alone does not answer it.
Is the vendor or any sub-processor subject to foreign law?Your processor agreement must say so, and sub-processors need your prior acceptance (Implementing Regulation, Article 17).
Does the task include sensitive data?Health data, biometric data used for identification and data revealing religious belief are Sensitive Data (PDPL, Article 1). Processing them requires a written impact assessment (Implementing Regulation, Article 25).
How fast will the vendor report a breach?You must notify SDAIA within 72 hours of becoming aware of a breach that may cause harm (Implementing Regulation, Article 24). The vendor's deadline has to be shorter than yours.
Is every action on a record logged?For health data, every processing stage must be documented along with the person responsible for it (Implementing Regulation, Article 26).

Our defaults are in-region hosting, on-prem or private-cloud deployment when data should stay in your environment, least-privilege access, audit logs on every task and contributors under NDA. If you need everything in-Kingdom (storage, processing and the people who open each record), write it into the scope and ask every vendor, us included, to confirm each location in writing. The full pre-annotation checklist is in PDPL and data residency: a checklist for MENA AI teams.

Sectors we support in the Kingdom

SectorTypical workRule to check first
Government and public sectorCitizen-service chat and voice in dialect, records digitization, and OCR on Hijri-dated forms (١٥/٠٣/١٤٤٧هـ)SDAIA data management standards for partners handling government data
Banking and financeKYC document extraction, complaint and call intent, transaction-narrative labelingSAMA's Cyber Security Framework 3.4.3: cloud services in Saudi Arabia "in principle", with explicit approval needed otherwise
HealthcareClinical transcription, triage intent in dialect, de-identificationHealth data is Sensitive Data, and Article 26 of the Implementing Regulation limits processing to the minimum necessary
TelecomCall transcription, churn and sentiment in dialect, voice-assistant dataRecordings and subscriber identifiers are personal data, so minimize them before labeling
Model buildersSFT, preference and evaluation data in Saudi varietiesUser prompts and logs used as seed data may contain personal data

We also work in legal, media and retail. See industries.

How a project runs

  1. Scoping. We agree the task, taxonomy, dialect mix, domain, modality, volume and residency terms, and which fields (names, national ID and Iqama numbers, phone numbers) are masked before anyone sees the data.
  2. Guidelines and gold set. We write guidelines covering Saudi-specific edge cases and build a gold set that linguists adjudicate. Contributors are vetted for the variety and domain, then trained and calibrated against the gold set.
  3. Pilot. A fixed-scope production run on your own data, in the same environment you will use for production.
  4. QA and adjudication. Hidden gold items check each annotator. Overlapping labels measure inter-annotator agreement, with the statistic chosen to suit the task (Cohen's kappa, Krippendorff's alpha, or span-level F1). Adjudicators resolve disagreements, and systematic errors go back into the guidelines and are reworked. The pilot ends with a quality report.
  5. Scale and delivery. Managed, embedded or enterprise engagements, delivered in your formats, with an audit trail on every task.

When to automate and when to use human annotators

Model pre-labeling makes sense only where you can show it is safe. For Saudi and Gulf data, that is often not the case yet. In our LabelBench audit, a local language model that scored 78% on binary Arabic toxicity fell to 25.3% exact-match accuracy on the real multi-label taxonomy and to 22.4% across five dialect regions (chance is 20%), on the fixed held-out manifest. It labeled 75% of Gulf examples as Egyptian. Our approach is to measure Safe Automation Coverage on your taxonomy and dialect mix and automate only the share that stays above your quality floor. Multi-label, dialect-sensitive and regulated-domain work stays with human annotators.

In-house team, crowd platform or specialist vendor?

In-house teamCrowd platformSpecialist Arabic vendor
Saudi dialect depthOnly the dialects you hireSelf-reported and hard to verifyScreened per variety
Data locationFully under your controlOften a global workforceSet by contract, including on-prem options
Domain expertsYou hire each oneRareLicensed professionals where needed
QA methodYou build itTools provided; you design the gold setsGold sets, agreement measurement and adjudication included
Best forStable, long-running core taxonomiesHigh-volume, low-risk, non-personal dataDialect-heavy, regulated or expert data

Start with a pilot

Send the task, target varieties, domain, modality, rough volume and residency constraints to hello@bayanatlabs.com or through our contact form. We come back with a fixed-scope pilot on your own data, with gold-standard QA and a quality report, usually within two weeks. You then decide whether to scale.

Frequently asked questions

Does using a Saudi-based annotation vendor make us PDPL-compliant?

No. A local vendor removes most cross-border transfer questions, but the controller obligations stay with you: a lawful basis for processing, an accurate privacy notice, minimization, impact assessments where required, and a processor agreement with the terms the Implementing Regulation lists. Treat vendor location as one control among several. This is general information, not legal advice.

Do we need Saudi annotators, or will any native Arabic speaker do?

It depends on the task. MSA documents, forms and OCR ground truth can be handled by a broader pool of native speakers with the right domain training. Dialect identification, sentiment, intent and speech transcription need annotators who speak the variety in your data: a non-Saudi speaker can read دحين and still miss that it marks a Hijazi speaker, or mis-transcribe fast Najdi call audio.

Can annotators work inside our own environment?

Yes. On-prem and private-cloud options let contributors work inside your infrastructure, so records never have to be exported to us. Otherwise we host in-region by default. In both setups access is least-privilege, every task has an audit log, and contributors work under NDA. Share your tooling, network and identity constraints at scoping so the pilot runs in the same setup as production.

How much does data annotation cost in Saudi Arabia?

We quote after scoping rather than publishing a rate card, because cost follows variables you control: modality (audio and video cost more per unit than short text), how rare the target varieties are, whether licensed professionals must do the labeling, how much overlap and adjudication the QA plan uses, and residency constraints such as on-prem work. A pilot on your own data gives you a cost basis before you commit to volume.

Is Bayanat Labs hiring data annotators in Saudi Arabia?

We recruit through our Talent Network: remote, flexible, paid-per-task Arabic data work that is free to apply for. Native speakers of Saudi varieties, linguists and licensed professionals such as physicians, lawyers and bankers are all relevant. Every contributor passes dialect and domain screening before working on client data.

Scope an in-region pilot

Tell us the task, the Saudi varieties, the domain and your residency constraints. We come back with a fixed-scope pilot, gold-standard QA and a quality report, usually within two weeks.

Talk to us