Checklist · Documents & OCR

How to evaluate an Arabic OCR dataset before you buy it: a 10-point audit

Ask for 50 to 100 pages, open the files yourself and score them. This checklist covers script direction, reading order, the three digit sets, tables, field-level accuracy, annotation consistency, splits and licensing, with the public resources to compare against.

Bayanat Labs Research· · 10 min read · اقرأ بالعربية

To evaluate an Arabic OCR dataset before you buy it, take your own sample of 50 to 100 pages, re-transcribe part of it with native readers, and score ten things: encoding, script direction and reading order, digits, letter rules, geometry, tables, fields, label agreement, splits and diversity, and provenance and license. Score each check separately, never as one blended accuracy number, because Arabic documents fail in specific places: a dot, a digit group, a reversed key-value pair, a merged table cell.

This checklist is for document-AI teams buying or licensing Arabic OCR, layout or structured-extraction data. It applies whether the dataset is a stocked collection, a public benchmark you plan to build on, or a custom collection a vendor will deliver. For the annotation rules themselves, see our Arabic OCR annotation page.

Key takeaways
  • Draw the sample yourself. Vendor-picked pages show the best case. Stratify by document type, handwriting share and capture quality.
  • Separate the levels. Text (CER), reading order, fields (exact match, F1), tables (TEDS) and answers (ANLS) fail independently and need their own scores.
  • Count templates, not pages. 10,000 pages from 12 templates (an illustrative case) is a small dataset for generalization. Ask for issuer and template IDs and splits that keep them apart.
  • Licensing is part of quality. Training, display and redistribution rights are separate, and public sets often mix licenses.

Why Arabic document datasets need their own audit

KITAB-Bench (Heakl et al., Findings of ACL 2025), an 8,809-sample Arabic OCR and document-understanding benchmark across 9 domains and 36 sub-domains, highlights complex fonts, numeral recognition errors, word elongation and table structure detection as remaining challenges. Its authors report that the best model reached only 65% accuracy on PDF-to-Markdown conversion. If models still fail there, the ground truth used to train and test them has to be right in exactly those places.

Three Arabic-specific facts drive most of the checks below:

  • Storage order differs from display order. Unicode stores text in logical (reading) order and the Unicode Bidirectional Algorithm reorders it for display. A line that mixes Arabic, Latin codes and digits can be copied in the wrong order and still look right on screen.
  • Several digit sets exist. Western 0–9, Arabic-Indic ٠–٩ (U+0660–0669) and Extended Arabic-Indic ۰–۹ (U+06F0–06F9) are different code points, described in chapter 9 (Middle East-I) of the Unicode Standard.
  • Dots and diacritics carry meaning. خبر ("news") and حبر ("ink") differ by one dot; درس ("he studied") and درّس ("he taught") by one shadda.

The 10-point Arabic OCR dataset audit

Run every check on the same sample. A red flag on checks 1, 2, 3, 9 or 10 usually ends the evaluation; the others can often be fixed by the vendor with a spec change and a re-delivery.

#CheckHow to inspectRed flag
1Unicode and encodingScan transcriptions for presentation-form code points (U+FB50–FDFF, U+FE70–FEFF), stray control characters and mixed normalization.Ligature glyphs such as ﻻ stored instead of ل + ا; text extracted from a PDF layer, not transcribed.
2Script direction and reading orderRender each line RTL beside its crop. Check mixed lines (Arabic + invoice codes, phone numbers, amounts) and the region order on multi-column pages.Reversed Arabic (ةروتافلا), digit groups in display order, columns read left to right.
3Digits and numeralsConfirm digits are stored as printed, with any normalized value in a separate field; check thousands and decimal separators.٤ silently converted to 4 in the ground truth; ١٬٢٥٠٫٠٠ stored as ١٢٥٠٠٠.
4Letter, dot and diacritic rulesRead the transcription guideline. Is it diplomatic (as written) or corrected (as intended)? Are harakat kept as printed? How are illegible spans marked?No written rule, or both conventions in one dataset; diacritics stripped in the ground truth.
5GeometryOverlay boxes or polygons on 20 pages. Word boxes must include dots and diacritics above and below the line.Boxes cut off dots; line polygons that merge two lines; coordinates for a different page size.
6TablesCheck cell boxes, row and column indices, merged cells and header flags on RTL tables.Columns indexed left to right with no stated convention; merged cells duplicated; structure and text scored as one number.
7Fields and key-value linksFor forms and invoices, check that each value is linked to its key, both the printed and normalized value are stored, and Hijri and Gregorian dates are labeled.Fields extracted by a model and never human-verified; totals that do not reconcile with line items.
8Annotation consistencyAsk for double-transcription and adjudication records; re-transcribe 10 to 20 pages yourself and compare.No agreement figures; disagreements resolved by "majority" with no senior reviewer; your re-transcription disagrees on digits or dots.
9Diversity and splitsAsk for counts of issuers, templates, writers and capture devices, and how test splits are separated.Only a page count; test pages share templates or writers with training pages.
10Provenance, consent and licenseAsk where every page came from, what personal data it contains and what the license permits: training, display, redistribution, derivatives.No written source permissions; real IDs or invoices without de-identification; a research-only source inside a commercial set.

Worked example: one invoice line, three ways to get it wrong

Take a line from a bilingual Saudi tax invoice. Saudi VAT Implementing Regulations (Article 53(5)) require invoice details in Arabic, with any other language as a translation, so Arabic, Latin and digits share lines.

PrintedDelivered ground truthWhich check fails
رقم الفاتورة: INV-2024-0183The Arabic key stored reversed, as some PDF text layers emit it1 and 2: extracted, not transcribed; logical order lost
الإجمالي: ١٬٢٥٠٫٠٠ ر.س١٢٥٠٠٠3: separators dropped, a different amount; no printed value kept
Total beside a handwritten correctionOnly the printed total, no link to the correction7: the field a validation model needs is missing

None of these errors shows up in a page-level accuracy figure. All three show up quickly with the RTL render-and-compare check.

Which metric for which level

Ask the vendor which of these they measured on their own ground truth, on what sample, and against whose second transcription. Then compute the same numbers on your sample.

LevelMetricArabic-specific detail
Line and word textCharacter and word error rate (CER, WER)Report with and without diacritics, and separately for digits and Latin runs. Normalize only at scoring time.
Reading orderOrder errors per page; region-order edit distanceScore mixed-direction lines on their own; they hide most bidi errors.
FieldsExact match on normalized values; precision, recall and F1 per field type; key-value link accuracyDates in both calendars, amounts with Arabic separators, ID and Iqama numbers.
TablesTEDS, the tree-edit-distance similarity introduced with PubTabNet (Zhong et al., 2019)State the column-order convention for RTL tables; score structure separately from cell text.
Document questionsANLS, from the ST-VQA paper (Biten et al., 2019) and used for DocVQAString similarity forgives small OCR slips but not a wrong digit in an amount; add exact match for numeric answers.
Label agreementInter-annotator agreement on a double-labeled sampleFor categorical labels, chance-corrected coefficients such as kappa or alpha; Artstein and Poesio (2008) survey which to use when.

Report every metric by slice: printed versus handwritten, document type, capture quality (flatbed scan, phone photo, fax). A dataset that scores well on clean scans and has no phone photos is not a dataset for a mobile onboarding flow.

Public Arabic OCR resources to compare against

Public sets are useful as a second opinion: run the same checks on a public sample and on the offered dataset, and compare. Check each license before any commercial use.

ResourceWhat it coversUse in an audit
KITAB-Bench (2025)8,809 samples, 9 domains: handwriting, tables, charts, PDF-to-Markdown and moreTask list and metrics to mirror; a reference for where models fail
KHATT1,000 handwritten forms from 1,000 writers, page, paragraph and line ground truthWriter-disjoint splitting; handwriting transcription conventions
Muharaf (NeurIPS 2024)Over 1,600 historic handwritten pages with line polygons, transcribed by archival expertsPolygon geometry on cursive text; a split license to learn from
FUNSD (2019, English)199 scanned forms with entities and linksA model for key-value linking labels; no Arabic
PubTabNet (2019, English)Table images with structured HTML; introduced TEDSTable scoring method

Two English sets are on the list on purpose: they define label formats and metrics that Arabic projects reuse. They say nothing about Arabic script, so they cannot replace an Arabic sample.

Ask for the dataset's documentation

Datasheets for Datasets (Gebru et al.) proposes documenting each dataset's motivation, composition, collection process and recommended uses. For an Arabic document set, the datasheet should answer at least:

  1. Document types and the count of issuers, templates and writers, separately from page count.
  2. Which pages are genuine, de-identified or synthetic, labeled per page.
  3. The transcription guideline version, the diplomatic-versus-corrected decision and the digit rule.
  4. Who transcribed and who adjudicated: native readers, screening, and the share double-transcribed.
  5. Split rules, and confirmation that no template, issuer or writer crosses train and test.
  6. Source permissions, personal-data handling and the license, with training, display and redistribution stated separately.

If personal data is involved (IDs, statements, medical forms), add residency and access questions from our PDPL and data residency checklist. For broader licensing questions on off-the-shelf data, see our guide to commercial Arabic datasets.

When no offered dataset passes

Often none will, because your documents are specific: one bank's statements, one ministry's forms, one retailer's receipts. Rendered synthetic pages can fill font coverage but not real capture conditions (see the limits of synthetic data). The practical answer is a custom collection or annotation of your own authorized pages, built against this checklist from the first page. The vendor questions that matter are in how to choose an Arabic data annotation vendor, and the labeling conventions in how to annotate Arabic text.

Why document-AI teams build Arabic document data with Bayanat Labs

When the dataset you need does not exist off the shelf, Bayanat Labs is the best partner to build it, because the hardest points in this audit, from template diversity and splits to source permissions and gold-standard QA, are built into how we deliver:

  • Built around your documents. Bayanat Documents is a custom scope: we start with one document type and agree source permissions, layout diversity, capture conditions and page count around the problem you need to solve.
  • The labels your model needs, and no more. Request transcription, coordinates and reading order on their own, or add layout, field extraction, document questions and validation checks.
  • Checks 9 and 10 built in. Genuine, de-identified and synthetic documents remain distinguishable, template diversity is reported separately from page count, and evaluation splits can separate issuers, templates, sources and writers.
  • Native readers with quality control. Vetted native speakers and linguists across 25+ Arabic varieties, screened before they touch client data, calibrated against gold pages with adjudication and rework loops, and an audit trail on every task (how our OCR annotation runs).
  • Sovereign by default. In-region hosting by default, on-prem and private-cloud options, least-privilege access and contributors under NDA.
  • Proof before scale. A fixed-scope pilot with gold-standard QA and a quality report, usually within two weeks.

We apply the same discipline to automation: our LabelBench audit shows how a high binary accuracy can hide failure on the real task, which is why we report by slice.

Frequently asked questions

How many pages do I need to audit an Arabic OCR dataset?

Enough to see every slice you care about. A practical start is 50 to 100 pages drawn by you, not the vendor, stratified by document type, print versus handwriting and capture quality, with at least a few pages from each claimed template or issuer. Re-transcribe 10 to 20 of them yourself to measure label quality, not just model quality.

What is a good character error rate for Arabic OCR ground truth?

There is no universal threshold, and a single blended number hides failures. Set targets per slice: clean printed text, handwriting, digits and Latin runs, phone photos. Compare the vendor's ground truth with your own double transcription on the same pages; disagreements on digits, dots and reading order matter more than a small overall CER.

Can I use public Arabic OCR datasets commercially?

Only if each dataset's own license allows it. Many research datasets restrict commercial use or redistribution. Muharaf, for example, distributes its public part under CC BY-NC-SA 4.0 and a restricted part under a proprietary license limited to research use, with no redistribution. Read the license file for every source, and keep training, display and redistribution rights separate.

Should synthetic rendered pages count toward dataset size?

Report them separately. Rendered pages are useful for fonts and rare characters, but they do not reproduce real scanners, stamps, folds or handwriting. Ask for genuine, de-identified and synthetic pages to be labeled as such, and evaluate only on genuine pages unless your production traffic is synthetic too.

What metric should I use for invoice and form extraction?

Score each field, not the page. Use exact match on normalized values (dates, amounts, ID numbers) and report precision, recall and F1 per field type, plus key-value linking accuracy. Keep a separate CER on the raw text, so you can tell an OCR miss from an extraction miss.

Need Arabic document data that passes this audit?

We build custom Arabic document datasets around your document types: transcriptions, reading order, fields, tables and evidence, with template diversity reported separately from page count and evaluation splits that can separate issuers, templates and writers.

Discuss your documents