Arabic Content Moderation Annotation: Toxicity Taxonomies That Work
A toxic/not-toxic label hides the errors that matter in Arabic. How to design a multi-label toxicity taxonomy, where dialect, sarcasm and religious language trip up annotators and models, and the QA and safety practices that keep moderation data trustworthy.
Arabic content moderation annotation, in short: it is the work of labeling Arabic posts, comments, messages and model outputs against a harm taxonomy, so that classifiers and safety filters can learn what to remove, restrict or escalate. It works when the taxonomy is multi-label and tied to policy actions, each item is labeled by a native speaker of its dialect, and guidelines spell out how to treat sarcasm, banter and religious language. A single toxic/not-toxic label hides most of the errors that matter.
Content note: this guide discusses offensive language but reproduces no slurs; examples use placeholders such as [SLUR] or [GROUP], following the presentation advice in Kirk et al. (2022).
Why do binary toxicity labels fail Arabic content moderation?
A moderation system does not make one decision. It removes some content, age-gates some and escalates some. A toxic/not-toxic label cannot drive those routes, and it flatters: when non-toxic items dominate a sample, a model can post a high binary score while getting the rare, severe categories wrong.
Arabic makes this worse. The same phrase can be banter in one dialect and an insult in another, religious vocabulary runs through everyday speech, and sarcasm borrows the language of praise. When annotators lack that context, the errors become training data. BSR's September 2022 human rights due diligence report on Meta found that, on a per-user basis, Arabic content saw greater over-enforcement than Hebrew content during May 2021. Possible causes it listed include content not being routed to reviewers who knew its dialect, and training data built from reviewers' decisions that "likely reproduces the errors of human reviewers due to lack of linguistic and cultural competence." It recommended routing Arabic content by dialect and region.
What LabelBench measured on Arabic toxicity
LabelBench, our reproducible audit of automated Arabic labeling, tests this gap. It uses AraTox, a human-annotated social-media dataset whose card lists Gulf, Levantine, Nile Basin, North African and Yemeni varieties plus MSA, with seven labels: appearance, cussing, hatred, racial, sexual, violence and NOT (non-toxic). LabelBench scores seeded samples from it two ways: collapsed to binary, and as the original multi-label task, where an item is correct only if the full set of labels matches.
On this fixed held-out manifest, a local language model (Qwen 2.5 3B) reached 78.0% on binary Arabic toxicity and 25.3% exact-match accuracy on the seven-label taxonomy. Fanar 1 9B, which is specialized for Arabic, improved both scores to 88.0% and 31.3%. Three details matter:
- Right number of labels, wrong categories. Qwen predicted 1.40 labels per item against 1.38 in the gold data, with per-label F1 from 10.3% on violence to 68.1% on NOT. The drop comes from confusing harm types, not from under-labeling.
- Specialization helps unevenly. Fanar reached 71.8% F1 on sexual but 20.0% on appearance.
- No generative run cleared the automation bar. Safe Automation Coverage (SAC) is our metric for the share of items a system can label automatically while a conservative accuracy bound stays above a quality floor. No evaluated generative system achieved non-zero SAC95 under the tied-confidence v0.1 protocol. Small supervised character-level diagnostics fared differently by task: 89.0% SAC95 on binary toxicity, but 0% on the seven-label task. SAC is a diagnostic, not a substitute for conformal risk control.
These are results for the evaluated local models, not a claim that all LLMs fail at Arabic. The narrower point: a binary benchmark says little about your production taxonomy.
How to design an Arabic toxicity taxonomy
Start from your policy, not a public dataset's label list: each label should map to a different action or reviewer. Make the scheme multi-label, because abuse routinely combines harm types: a sexualized insult aimed at a nationality is both sexual and identity-based hate. The AraTox categories are a sensible base layer:
| Label | Definition for the guideline | Pattern (placeholders) | Rule that prevents most disagreements |
|---|---|---|---|
| Cussing | Profanity or vulgar insult aimed at no protected attribute | يا [شتيمة] (ya [INSULT], "you [INSULT]") | Label profanity even when friendly; record banter as an attribute |
| Sexual | Sexual content, innuendo or sexualized insult | [SEXUAL INSULT] + a person's name or a family relation | Insults involving female relatives count here even when the target is a man |
| Hatred | Attack or dehumanization based on religion, sect, gender, nationality or other group identity | كل [الفئة] … (kull [al-fi'a], "all [GROUP] are …") | Needs a group, or a person attacked for membership; a general curse is not hatred |
| Racial | Attack based on race, skin color, ethnicity or origin | [ETHNIC SLUR] used as a term of address | Kept apart from hatred so race-based abuse is measurable |
| Appearance | Mockery of body, face, weight, disability or looks | شكلك [وصف] (shaklak [DESCRIPTION], "you look [DESCRIPTION]") | Add hatred for disability mockery if your policy protects disability |
| Violence | Threat, incitement or glorification of physical harm | "I will [HARM] you" or "someone should [HARM] [GROUP]" | Grade idiomatic hyperbole with severity; don't drop the label |
| NOT | No toxic content, including reporting, quoting or condemning abuse | A news report that quotes an insult | Counter-speech and reporting are NOT, with a "quotes harm" flag |
Then add attributes, not labels. Liebeskind et al. (2025) propose a layered Arabic taxonomy covering the target (individual, group, or individual attacked through group stereotypes), whether the target is present, vulgarity, severity (insult, hate speech, discrediting, threat) and explicit versus implicit offence. For production, five attributes earn their cost: target type, severity, explicit or implicit, context needed (not judgeable without the parent post) and dialect, so quality can be measured per variety.
Dialect, sarcasm and religious references: where Arabic toxicity labels go wrong
Dialect and script
Insult strength is regional. Calling someone a donkey, يا حمار (ya ḥmār), is an insult across the Arabic-speaking world, yet between friends it is often banter. Curses phrased as prayers, such as الله ياخذك (Allah yākhdhak, "may God take you"), range from exasperated joke to real hostility. The same content also arrives in Arabizi (ya 7mar, with 7 for ح) or with deliberate misspellings to dodge keyword filters, so guidelines need examples in every script your platform sees. Our guide to MSA versus dialectal Arabic explains why variety coverage sets a model's ceiling.
Sarcasm
Sarcastic abuse often borrows praise: ما شاء الله عليك، عبقري (mā shā' Allāh ʿalēk, ʿabqari, "mashallah, what a genius") under a post that got something wrong. The ArSarcasm authors re-annotated Arabic sentiment corpora and found 16% of 10,547 tweets sarcastic. They also note the task's subjectivity and label variation from annotator bias. The rule: label the intended meaning, set a sarcasm flag, and send items whose meaning depends on the thread to "context needed" instead of forcing a guess.
Religious references
Religious language causes costly errors in both directions. Formulas such as حسبي الله ونعم الوكيل (ḥasbiya Allāh wa-niʿma al-wakīl, "God is sufficient for me") express grievance and are not toxic in themselves, even when aimed at someone. The clearest public example is شهيد (shahīd). Meta estimated that the word and its variants accounted for more content removals than any other single word or phrase on its platforms, and in March 2024 the Oversight Board advised ending its blanket removal. The Board noted that the word also describes people who die serving their country or as victims of violence or disaster, and is used as a name, and it recommended removal only alongside specific signals of violence.
The opposite failure is real too. Calling a specific person or group an unbeliever (كافر, kāfir) is descriptive in a theology discussion and can work as incitement in a sectarian thread. Albadi et al. (2018) built one of the first public Arabic datasets for religious hate speech from tweets about religious groups, a sign of how much of this category turns on who is targeted rather than on vocabulary. The principle for all three cases: label the target and the intent, never the keyword.
Why Arabic moderation needs dialect-matched annotators
General fluency is not enough: a reviewer from one region can over-read an insult that is mild elsewhere or miss a local coded reference, one of the possible root causes BSR identified. In practice:
- Route by region to annotators screened for that dialect, using metadata or a dialect-ID pass, and re-route when the guess is wrong.
- Make "dialect unfamiliar" a valid answer. A skipped item is better data than a guess.
- Adjudicate religious, political and ethnic items across regions, and ship annotator region and adjudication history in the metadata so any label can be audited.
Arabic content moderation annotation: guidelines and QA
A moderation guideline is a decision procedure, not a glossary. Our Arabic annotation guidelines template gives the structure; for toxicity, these checks matter most:
- Decide in a fixed order: target, intent, labels, then severity, so annotators do not anchor on a vulgar word.
- Give every label positive, negative and near-miss examples in several dialects and Arabizi.
- Stratify the gold set by label and dialect, oversampling rare labels such as violence (Qwen's weakest label in LabelBench).
- Report agreement per label and per dialect. Overall agreement on a mostly-NOT sample can hide disagreement on threats, so compute kappa or alpha separately for each label and each dialect.
- Score exact sets as well as per-label F1, and adjudicate with written rationales that become new guideline examples.
- Measure before you automate: estimate per label and dialect how much can leave the human queue at your quality floor, as SAC does. Our guide to LLM-as-a-judge in Arabic covers the calibration protocol, and how to annotate Arabic text covers normalization and span rules.
How to protect annotators handling harmful Arabic content
Moderation work exposes annotators to abuse and threats for hours, in their own language. Kirk et al. argue that this exposure is unavoidable but can be limited. Their recommendations, plus standard operating practice, translate into workflow:
- Limit exposure: active learning so fewer items need labels, harmful words masked by default with a reveal control, dummy data while building pipelines.
- Make it opt-in: say which categories annotators will see, and let them leave one without losing other work.
- Support and debrief: mental health support, breaks and rotation, and debriefs that feed the guidelines.
- Escalate the worst material: a written path for possibly illegal content, so no annotator decides alone.
At Bayanat Labs, moderation data follows the same Vet → Train → Produce → QA → Deliver loop as our other Arabic data annotation services: contributors pass dialect and domain screening before touching client data, work is calibrated against gold standards with adjudication, and every task has an audit trail. When the content is model output rather than user posts, the same taxonomy feeds Arabic AI red teaming, where severity grading matters most. Comparing providers? Use our vendor RFP checklist.
- Binary scores flatter. On LabelBench's fixed held-out manifest, a local model fell from 78.0% on binary Arabic toxicity to 25.3% on the seven-label taxonomy.
- Design labels around actions. Use a multi-label scheme with one label per policy action, plus target, severity, explicitness and context-needed attributes.
- Label intent, not keywords. Dialect banter, sarcasm and religious vocabulary drive errors in both directions.
- Match annotators to dialects. Route by region, allow "dialect unfamiliar", and adjudicate culturally loaded items across regions.
- Measure per label and per dialect. Overall agreement hides disagreement on rare, severe categories, and those are the ones that matter.
- Protect the annotators. Masking, opt-in, rotation, support and escalation are part of data quality.
Frequently asked questions
Can an LLM moderate Arabic content without human review?
Not safely by default. In LabelBench v0.1, no evaluated generative system achieved non-zero Safe Automation Coverage at a 95% quality floor under the tied-confidence protocol, even where binary accuracy looked respectable. That is a diagnostic of the evaluated local models and prompts, not proof that no model can help. The workable pattern is model pre-labeling with confidence routing, measured per label and per dialect against a native-adjudicated gold set.
How many labels should an Arabic toxicity taxonomy have?
As many as your policy has distinct actions, and no more. If profanity is age-gated, threats are escalated and identity-based hate is removed, those need separate labels. Seven categories, as in AraTox, is a reasonable starting size. Every extra label splits your examples further and adds another boundary for annotators to disagree on, so rare labels quickly become too thin to measure. Put nuance into attributes such as target and severity instead of adding labels.
Are there public Arabic toxicity and hate speech datasets?
Yes. Examples include AraTox (seven labels, multi-dialect, CC BY 4.0), the OSACT4 shared-task data on offensive language and hate speech, Albadi et al.'s religious hate speech tweets and L-HSAB for Levantine. Their label definitions differ, so map each to your own taxonomy before merging, and check every license for commercial use.
Do Arabic moderators need to speak the specific dialect they review?
For anything beyond obvious profanity, yes. Insults, banter, imprecations and coded references are regional, and a fluent speaker from elsewhere can misjudge severity or miss the target. Route by region, let annotators flag 'dialect unfamiliar' without penalty, and adjudicate religious or political items across regions.