Arabic Annotation Guidelines: A Free Template for Dialect, Diacritics & Arabizi
Generic guideline templates stop where Arabic gets hard: dialect tags, diacritics, clitics, Arabizi, sarcasm and religious formulas. This one starts there. Copy the 12 sections, fill them in for your task, and calibrate them with a pilot before you scale.
Arabic annotation guidelines are the written rules that make annotators label Arabic data the same way every time: what each label means, which dialect, register and script tags apply, how to treat diacritics, clitics and Arabizi, and how to rule on sarcasm, negation and religious expressions. Generic templates cover definitions and examples, then stop. The free template below has 12 sections to copy, with Arabic examples, decision trees and a pilot plan.
Most public Arabic guidelines were written for a single research corpus. This page turns what they share into a structure for your own task. It follows common practice and the methodology we apply at Bayanat Labs (gold standards, adjudication, rework loops), and builds on our playbook for annotating Arabic text.
What must Arabic annotation guidelines cover that English ones don't?
The generic advice still holds: explain the task, define labels, show examples, pilot. Arabic adds decisions that change the label itself:
| Area | Why it is specific to Arabic | Template section |
|---|---|---|
| Diglossia and dialect | MSA and dialect mix inside one item; a word can mean different things across varieties | 6, 9 |
| Orthography | Hamza, taa marbuta and alef maqsura vary; diacritics are usually absent | 5 |
| Clitics | Conjunctions, prepositions, the article and pronouns attach to words | 4 |
| Script | Arabizi and Latin-script English or French inside Arabic lines | 7 |
| Negation | Negators differ across dialects (مش, مو, ما…ش) | 8 |
| Religious formulas | Very frequent phrases whose sentiment depends on context | 8 |
Even a correction task needs a dialect policy. The QALB error-annotation guidelines (Zaghouani et al., LREC 2014) include dialectal word usage among seven error categories, split dialectal issues into five types and correct only three. The authors report that several guideline changes were needed on how to handle dialect words, and on whether to correct or ignore certain word categories.
The Arabic annotation guidelines template: 12 sections to copy
Copy the list and fill in each prompt. It works for classification (sentiment, intent, toxicity) and span tasks alike.
- Purpose and downstream use. What the model does with these labels, who acts on its output, and which error costs more. Annotators rule better on ambiguous items when they know what the label is for.
- Data, varieties and scripts in scope. Sources, time window, expected varieties and registers with rough shares, scripts, and what to do with out-of-scope items (flag, don't label). Warn annotators here if the data may contain disturbing material.
- Label definitions. Per label: one-sentence definition, what it includes, what it excludes, its most confusable neighbour and the tie-break rule. Keep label IDs stable; add the Arabic name annotators see.
- Unit and span boundaries. Item, sentence or span; character offsets on raw text; whether attached و, ف, ب, ل sit inside spans; the article, titles and honorifics; nesting.
- Text handling and diacritics. The raw text is never edited: no spelling fixes, no added or removed diacritics. Record what the clean-up script changed, with its version, and what to do when an undiacritized word has two readings with different labels.
- Register, dialect and script tags. Fields, allowed values (including "mixed" and "undetermined"), level of detail and the evidence rule.
- Code-switching and Arabizi. Label text as written; never transliterate inside a labeling task. For token-level language tags, reuse the labels of the first code-switching shared task (Solorio et al., 2014), which covered MSA–dialect and has a "mixed" label for hybrid words. For anything annotators type, name a spelling standard such as CODA*, a conventional orthography for dialects.
- Pragmatics. One rule each for sarcasm, negation, religious and social formulas, quoted speech, questions and wishes.
- Examples. Per label: a clear positive, a near-miss from the most confusable label, and a borderline item with its ruling, spread across every variety and script, each with a translation and a "why".
- Edge cases and escalation. Decision trees, plus flags that pull an item from the queue: cannot judge, needs a specialist, personal data visible, harmful content. Name who answers questions, and how fast.
- Quality protocol. Hidden gold and its frequency, share of double-labeled items, the agreement measure and how its target was set, adjudicators per variety, rework triggers, feedback to annotators.
- Versioning and change log. A version on the document and on every label, one log row per ruling, and which language copy is authoritative.
How should dialect and register tags work?
Tag each item, not the dataset: register (MSA, dialectal, mixed), variety and script, at the coarsest level the product needs. The ArSarcasm dataset (Abu Farha and Magdy, 2020) used five values: Egyptian, Gulf, Levantine, Maghrebi and MSA. Then:
- Tag from linguistic evidence only. A dialect word, negator or verb prefix counts; the topic, a place name or the author's profile does not.
- "Undetermined" is a correct answer for short or formulaic items. A forced guess is noise.
- Route by variety. Annotators tag and adjudicate only varieties they speak natively.
Automation fails here most visibly. On LabelBench v0.1, Qwen 2.5 3B, a local language model, reached 22.4% exact-match accuracy on five dialect regions on the fixed held-out manifest (chance 20%). Our analysis of MSA versus dialect coverage covers which varieties to include.
Normalization, diacritics and span-boundary rules
Write these as rule cards: the rule, an example that follows it, one that breaks it. The playbook has the full normalization table; the core rules:
- Offsets refer to raw text. Normalized views are derived by a versioned script, never typed by annotators.
- Diacritics stay as written. When علم could be ʿilm ("knowledge") or ʿalam ("flag"), annotators resolve it from context or mark "ambiguous". They do not add marks.
- Attached conjunctions and prepositions sit outside spans. In وبيروت (wa-Bayrūt, "and Beirut") the entity is بيروت. In للرياض ("to Riyadh") the span is لرياض, with الرياض recorded as the canonical form.
Titles, nested organizations and names that are also common words are covered in our guide to Arabic NER annotation.
Sarcasm, negation and religious expressions in Arabic labeling guidelines
Sarcasm. Keep a sarcasm flag separate from sentiment, and label sentiment by intended meaning. ArSarcasm defines sarcasm as an utterance expressing ridicule where the intended meaning differs from the apparent one. Re-annotating existing Arabic sentiment datasets, its authors found 16% of 10,547 tweets sarcastic, and more than half of the tweets originally labeled positive changed label; in some of their examples the original annotator had missed the sarcasm.
Negation. List the negators of every variety in scope and state the scope rule: negation flips the evaluation it attaches to, not the whole item. Egyptian wraps the verb (ماعجبنيش, "I didn't like it"); Egyptian and Levantine put مش before adjectives; Gulf speakers often use مو. Then rule on negated negatives: "not bad" can be mild praise or faint criticism. Pick one reading and give an example.
Religious and social formulas. The default: a formula carries no sentiment of its own; the rest of the item decides.
| Expression | Literal meaning | Common uses | Default rule |
|---|---|---|---|
| إن شاء الله in shāʾ allāh | "If God wills" | A sincere plan, a polite hedge, a soft refusal | Neutral alone; label what surrounds it |
| الحمد لله al-ḥamdu lillāh | "Praise be to God" | Gratitude, but also patience after bad news: الحمد لله على كل حال ("in every circumstance") | Not automatically positive; read the event it answers |
| ما شاء الله mā shāʾ allāh | "What God has willed" | Admiration, often to ward off envy; can be sarcastic | Positive only if the content is; check sarcasm |
| حسبي الله ونعم الوكيل ḥasbiya llāhu wa-niʿma l-wakīl | "God suffices me, and He is the best guardian" | Grievance at being wronged | Negative; not an insult in toxicity tasks |
For toxicity tasks, add the matching rule: religious language is not offensive in itself; attacking people for their religion is. AHA-Memes (Kmainasi et al., 2026) lists religion among protected characteristics and used bilingual English and Arabic guidelines to capture dialectal and cultural cues. Write these rules before annotation starts: the authors of the ThatiAR subjectivity dataset (Suwaileh et al., 2024) found annotators strongly influenced by their political, cultural and religious backgrounds, especially early on.
Writing positive and negative examples in Arabic
Examples do more work than definitions. Draw them from real traffic, not translated English, strip personal data, and cover every variety and script. The rows illustrate a customer-feedback sentiment task (Positive, Negative, Neutral, Mixed, plus a sarcasm flag); the sentences are illustrative, written for this guide.
| Label | Example | Translation | Why |
|---|---|---|---|
| Positive | الخدمة مرة زينة al-khidma marra zēna | "The service is really good." | Saudi: مرة means "very" here, not "once". |
| Negative | ماعجبنيش التحديث الجديد maʿagabnīsh et-taḥdīs el-gedīd | "I didn't like the new update." | Egyptian negation wraps the verb; a keyword match on عجب ("like") scores it positive. |
| Negative + sarcasm | والله خدمة ممتازة، ثلاث ساعات انتظار wallāh khidma mumtāza, thalāth sāʿāt intiẓār | "Wow, excellent service: three hours of waiting." | The complaint contradicts the praise. Label the intended meaning; flag the sarcasm. |
| Neutral | إن شاء الله أجرب التطبيق بكرة in shāʾ allāh ajarrib at-taṭbīq bukra | "God willing, I'll try the app tomorrow." | A plan, not an evaluation. A near-miss for Positive. |
| Positive (mild) | التطبيق الجديد مش بطال et-taṭbīʾ el-gedīd mish baṭṭāl | "The new app isn't bad." | Negated negative; this template rules it mild positive. Whatever you choose, write it down. |
| Mixed | el app 7elw bas bate2 | "The app is nice but slow." | Arabizi (7 for ح, 2 for hamza); two evaluations joined by bas, "but". Tag the script; don't transliterate. |
| Positive | ما شاء الله، التحويل وصل بثواني mā shāʾ allāh, at-taḥwīl wiṣil bi-thawānī | "Mashallah, the transfer arrived in seconds." | Positive because of the fast transfer, not the formula. Compare the Neutral row. |
Edge-case decision trees for Arabic labeling guidelines
Trees fix the order of questions, which is where much disagreement starts. Keep each on one screen.
Tree A: sentiment.
- Arabic (any script) and intelligible? If not, flag "out of scope" and stop.
- Personal data or harmful content that section 10 routes elsewhere? Flag and stop.
- Set aside religious and social formulas. Does the rest evaluate anything? If not, label Neutral.
- Is the literal evaluation contradicted by context? Set the sarcasm flag; label the intended meaning.
- Is the evaluation negated? Apply the negation rule, including negated negatives.
- Evaluations in both directions? Mixed. Otherwise Positive or Negative.
Tree B: dialect tag.
- Any dialect marker? If not and the grammar is standard, tag MSA; if too short to tell, "undetermined".
- Markers from one variety? Tag it. From several? Tag the dominant one, register "mixed".
- Not a variety you speak natively? Skip so it routes to a native speaker.
Tree C: stop and ask. Flag rather than guess when the label needs specialist knowledge (medical, legal, religious), when an undiacritized word has two readings with different labels, or when two rules conflict. Each answer becomes a new example in section 9.
How do you version annotation guidelines during a project?
Every rule change splits your dataset into "before" and "after" unless you track it:
- Two-part versions. Bump the major number when a definition, span rule or tag value changes; the minor number for new examples and clarifications.
- Stamp every label with its guideline version, so rework targets only affected items.
- One change-log row per ruling: date, version, section, old and new rule, triggering example, and the rework decision (none, spot-check, relabel matches).
- Release on a schedule once production starts, with bilingual copies updated together and before-and-after examples for annotators.
How to calibrate Arabic annotation guidelines with a pilot
Published Arabic projects plan for revision. QALB annotated a trial sample of each text type (50 documents for news comments) and used the trials to revise its guidelines. The BAREC readability guidelines (Habash et al., 2025) were refined through iterative training with native Arabic-speaking educators, reaching a quadratic weighted kappa of 81.8% in the final annotation phase. For your pilot:
- Stratify the sample across every variety, script and source in section 2, weighted toward hard cases.
- Overlap fully: several annotators per variety label the same items independently.
- Seed hidden gold. ArSarcasm used 100 hidden test questions and dropped annotators, and all their labels, below 80% on them.
- Sort disagreements by cause: unclear rule, missing rule, genuine ambiguity, annotator error, tool problem. Only the first two change the guideline.
- Rule, version, repeat on fresh items until most remaining disagreement is genuine ambiguity.
- Set targets per task and variety from the pilot. ArSarcasm's annotators agreed 80.7% of the time on sentiment and 89.3% on sarcasm, on the same tweets.
Pre-labeling with a model? Test your taxonomy on it in the pilot. On LabelBench v0.1, Qwen 2.5 3B, a local language model, reached 78% exact-match accuracy on binary Arabic toxicity but 25.3% on the seven-label taxonomy, on the fixed held-out manifest. Our complete guide to Arabic data labeling sets out when pre-labels are safe.
At Bayanat Labs, a fixed-scope pilot is where a guideline meets real data. Contributors pass dialect and domain screening before touching client data, work is calibrated against gold standards with adjudication and rework loops, and every task has an audit trail. Our Arabic data annotation services cover NER and spans, intent and sentiment, transcription and segmentation; for speech, see our Arabic transcription guidelines.
- Generic templates stop too early. Arabic needs written rules for dialect, script, normalization, clitics, negation and formulas.
- Formulas carry no sentiment alone. Label what surrounds إن شاء الله or الحمد لله; keep sarcasm as a separate flag.
- Every example needs a "why" and a near-miss, in each variety and script.
- Version the guideline, stamp the labels, pilot until remaining disagreement is genuine ambiguity.
Frequently asked questions
Should Arabic annotation guidelines be written in Arabic or English?
Write the rules in the language your annotators think in, and keep every example in its original Arabic. Many teams keep a bilingual document: English for the ML team, Arabic for annotators. If you do, name one copy as authoritative and ship every change in both languages in the same release.
How long should annotation guidelines be?
As short as the task allows. Keep a core document annotators can read in one sitting, and move examples and decision trees to an appendix they search while working. The document grows with every pilot ruling. When a section passes a few pages, split it into a quick-reference card and a detailed annex.
Who should write Arabic annotation guidelines?
A pair: the ML or product owner, who knows what the model needs, and a senior Arabic linguist who knows the varieties in scope. The first sets purpose and label boundaries; the second writes the dialect, orthography and pragmatics rules and picks the examples. Senior adjudicators should propose rulings, because they see failures first.
How often should annotation guidelines change during a project?
Often during the pilot, rarely afterwards. Expect several revisions while calibrating, then batch changes into scheduled releases once production starts, so nobody relearns rules mid-batch. New examples can ship quickly. A changed label definition needs a recorded decision about which labeled items to rework.