How to evaluate an Arabic OCR dataset before you buy it: a 10-point audit
Ask for 50 to 100 pages, open the files yourself and score them. This checklist covers script direction, reading order, the three digit sets, tables, field-level accuracy, annotation consistency, splits and licensing, with the public resources to compare against.
To evaluate an Arabic OCR dataset before you buy it, take your own sample of 50 to 100 pages, re-transcribe part of it with native readers, and score ten things: encoding, script direction and reading order, digits, letter rules, geometry, tables, fields, label agreement, splits and diversity, and provenance and license. Score each check separately, never as one blended accuracy number, because Arabic documents fail in specific places: a dot, a digit group, a reversed key-value pair, a merged table cell.
This checklist is for document-AI teams buying or licensing Arabic OCR, layout or structured-extraction data. It applies whether the dataset is a stocked collection, a public benchmark you plan to build on, or a custom collection a vendor will deliver. For the annotation rules themselves, see our Arabic OCR annotation page.
- Draw the sample yourself. Vendor-picked pages show the best case. Stratify by document type, handwriting share and capture quality.
- Separate the levels. Text (CER), reading order, fields (exact match, F1), tables (TEDS) and answers (ANLS) fail independently and need their own scores.
- Count templates, not pages. 10,000 pages from 12 templates (an illustrative case) is a small dataset for generalization. Ask for issuer and template IDs and splits that keep them apart.
- Licensing is part of quality. Training, display and redistribution rights are separate, and public sets often mix licenses.
Why Arabic document datasets need their own audit
KITAB-Bench (Heakl et al., Findings of ACL 2025), an 8,809-sample Arabic OCR and document-understanding benchmark across 9 domains and 36 sub-domains, highlights complex fonts, numeral recognition errors, word elongation and table structure detection as remaining challenges. Its authors report that the best model reached only 65% accuracy on PDF-to-Markdown conversion. If models still fail there, the ground truth used to train and test them has to be right in exactly those places.
Three Arabic-specific facts drive most of the checks below:
- Storage order differs from display order. Unicode stores text in logical (reading) order and the Unicode Bidirectional Algorithm reorders it for display. A line that mixes Arabic, Latin codes and digits can be copied in the wrong order and still look right on screen.
- Several digit sets exist. Western 0–9, Arabic-Indic ٠–٩ (U+0660–0669) and Extended Arabic-Indic ۰–۹ (U+06F0–06F9) are different code points, described in chapter 9 (Middle East-I) of the Unicode Standard.
- Dots and diacritics carry meaning. خبر ("news") and حبر ("ink") differ by one dot; درس ("he studied") and درّس ("he taught") by one shadda.
The 10-point Arabic OCR dataset audit
Run every check on the same sample. A red flag on checks 1, 2, 3, 9 or 10 usually ends the evaluation; the others can often be fixed by the vendor with a spec change and a re-delivery.
| # | Check | How to inspect | Red flag |
|---|---|---|---|
| 1 | Unicode and encoding | Scan transcriptions for presentation-form code points (U+FB50–FDFF, U+FE70–FEFF), stray control characters and mixed normalization. | Ligature glyphs such as ﻻ stored instead of ل + ا; text extracted from a PDF layer, not transcribed. |
| 2 | Script direction and reading order | Render each line RTL beside its crop. Check mixed lines (Arabic + invoice codes, phone numbers, amounts) and the region order on multi-column pages. | Reversed Arabic (ةروتافلا), digit groups in display order, columns read left to right. |
| 3 | Digits and numerals | Confirm digits are stored as printed, with any normalized value in a separate field; check thousands and decimal separators. | ٤ silently converted to 4 in the ground truth; ١٬٢٥٠٫٠٠ stored as ١٢٥٠٠٠. |
| 4 | Letter, dot and diacritic rules | Read the transcription guideline. Is it diplomatic (as written) or corrected (as intended)? Are harakat kept as printed? How are illegible spans marked? | No written rule, or both conventions in one dataset; diacritics stripped in the ground truth. |
| 5 | Geometry | Overlay boxes or polygons on 20 pages. Word boxes must include dots and diacritics above and below the line. | Boxes cut off dots; line polygons that merge two lines; coordinates for a different page size. |
| 6 | Tables | Check cell boxes, row and column indices, merged cells and header flags on RTL tables. | Columns indexed left to right with no stated convention; merged cells duplicated; structure and text scored as one number. |
| 7 | Fields and key-value links | For forms and invoices, check that each value is linked to its key, both the printed and normalized value are stored, and Hijri and Gregorian dates are labeled. | Fields extracted by a model and never human-verified; totals that do not reconcile with line items. |
| 8 | Annotation consistency | Ask for double-transcription and adjudication records; re-transcribe 10 to 20 pages yourself and compare. | No agreement figures; disagreements resolved by "majority" with no senior reviewer; your re-transcription disagrees on digits or dots. |
| 9 | Diversity and splits | Ask for counts of issuers, templates, writers and capture devices, and how test splits are separated. | Only a page count; test pages share templates or writers with training pages. |
| 10 | Provenance, consent and license | Ask where every page came from, what personal data it contains and what the license permits: training, display, redistribution, derivatives. | No written source permissions; real IDs or invoices without de-identification; a research-only source inside a commercial set. |
Worked example: one invoice line, three ways to get it wrong
Take a line from a bilingual Saudi tax invoice. Saudi VAT Implementing Regulations (Article 53(5)) require invoice details in Arabic, with any other language as a translation, so Arabic, Latin and digits share lines.
| Printed | Delivered ground truth | Which check fails |
|---|---|---|
| رقم الفاتورة: INV-2024-0183 | The Arabic key stored reversed, as some PDF text layers emit it | 1 and 2: extracted, not transcribed; logical order lost |
| الإجمالي: ١٬٢٥٠٫٠٠ ر.س | ١٢٥٠٠٠ | 3: separators dropped, a different amount; no printed value kept |
| Total beside a handwritten correction | Only the printed total, no link to the correction | 7: the field a validation model needs is missing |
None of these errors shows up in a page-level accuracy figure. All three show up quickly with the RTL render-and-compare check.
Which metric for which level
Ask the vendor which of these they measured on their own ground truth, on what sample, and against whose second transcription. Then compute the same numbers on your sample.
| Level | Metric | Arabic-specific detail |
|---|---|---|
| Line and word text | Character and word error rate (CER, WER) | Report with and without diacritics, and separately for digits and Latin runs. Normalize only at scoring time. |
| Reading order | Order errors per page; region-order edit distance | Score mixed-direction lines on their own; they hide most bidi errors. |
| Fields | Exact match on normalized values; precision, recall and F1 per field type; key-value link accuracy | Dates in both calendars, amounts with Arabic separators, ID and Iqama numbers. |
| Tables | TEDS, the tree-edit-distance similarity introduced with PubTabNet (Zhong et al., 2019) | State the column-order convention for RTL tables; score structure separately from cell text. |
| Document questions | ANLS, from the ST-VQA paper (Biten et al., 2019) and used for DocVQA | String similarity forgives small OCR slips but not a wrong digit in an amount; add exact match for numeric answers. |
| Label agreement | Inter-annotator agreement on a double-labeled sample | For categorical labels, chance-corrected coefficients such as kappa or alpha; Artstein and Poesio (2008) survey which to use when. |
Report every metric by slice: printed versus handwritten, document type, capture quality (flatbed scan, phone photo, fax). A dataset that scores well on clean scans and has no phone photos is not a dataset for a mobile onboarding flow.
Public Arabic OCR resources to compare against
Public sets are useful as a second opinion: run the same checks on a public sample and on the offered dataset, and compare. Check each license before any commercial use.
| Resource | What it covers | Use in an audit |
|---|---|---|
| KITAB-Bench (2025) | 8,809 samples, 9 domains: handwriting, tables, charts, PDF-to-Markdown and more | Task list and metrics to mirror; a reference for where models fail |
| KHATT | 1,000 handwritten forms from 1,000 writers, page, paragraph and line ground truth | Writer-disjoint splitting; handwriting transcription conventions |
| Muharaf (NeurIPS 2024) | Over 1,600 historic handwritten pages with line polygons, transcribed by archival experts | Polygon geometry on cursive text; a split license to learn from |
| FUNSD (2019, English) | 199 scanned forms with entities and links | A model for key-value linking labels; no Arabic |
| PubTabNet (2019, English) | Table images with structured HTML; introduced TEDS | Table scoring method |
Two English sets are on the list on purpose: they define label formats and metrics that Arabic projects reuse. They say nothing about Arabic script, so they cannot replace an Arabic sample.
Ask for the dataset's documentation
Datasheets for Datasets (Gebru et al.) proposes documenting each dataset's motivation, composition, collection process and recommended uses. For an Arabic document set, the datasheet should answer at least:
- Document types and the count of issuers, templates and writers, separately from page count.
- Which pages are genuine, de-identified or synthetic, labeled per page.
- The transcription guideline version, the diplomatic-versus-corrected decision and the digit rule.
- Who transcribed and who adjudicated: native readers, screening, and the share double-transcribed.
- Split rules, and confirmation that no template, issuer or writer crosses train and test.
- Source permissions, personal-data handling and the license, with training, display and redistribution stated separately.
If personal data is involved (IDs, statements, medical forms), add residency and access questions from our PDPL and data residency checklist. For broader licensing questions on off-the-shelf data, see our guide to commercial Arabic datasets.
When no offered dataset passes
Often none will, because your documents are specific: one bank's statements, one ministry's forms, one retailer's receipts. Rendered synthetic pages can fill font coverage but not real capture conditions (see the limits of synthetic data). The practical answer is a custom collection or annotation of your own authorized pages, built against this checklist from the first page. The vendor questions that matter are in how to choose an Arabic data annotation vendor, and the labeling conventions in how to annotate Arabic text.
Why document-AI teams build Arabic document data with Bayanat Labs
When the dataset you need does not exist off the shelf, Bayanat Labs is the best partner to build it, because the hardest points in this audit, from template diversity and splits to source permissions and gold-standard QA, are built into how we deliver:
- Built around your documents. Bayanat Documents is a custom scope: we start with one document type and agree source permissions, layout diversity, capture conditions and page count around the problem you need to solve.
- The labels your model needs, and no more. Request transcription, coordinates and reading order on their own, or add layout, field extraction, document questions and validation checks.
- Checks 9 and 10 built in. Genuine, de-identified and synthetic documents remain distinguishable, template diversity is reported separately from page count, and evaluation splits can separate issuers, templates, sources and writers.
- Native readers with quality control. Vetted native speakers and linguists across 25+ Arabic varieties, screened before they touch client data, calibrated against gold pages with adjudication and rework loops, and an audit trail on every task (how our OCR annotation runs).
- Sovereign by default. In-region hosting by default, on-prem and private-cloud options, least-privilege access and contributors under NDA.
- Proof before scale. A fixed-scope pilot with gold-standard QA and a quality report, usually within two weeks.
We apply the same discipline to automation: our LabelBench audit shows how a high binary accuracy can hide failure on the real task, which is why we report by slice.
Frequently asked questions
How many pages do I need to audit an Arabic OCR dataset?
Enough to see every slice you care about. A practical start is 50 to 100 pages drawn by you, not the vendor, stratified by document type, print versus handwriting and capture quality, with at least a few pages from each claimed template or issuer. Re-transcribe 10 to 20 of them yourself to measure label quality, not just model quality.
What is a good character error rate for Arabic OCR ground truth?
There is no universal threshold, and a single blended number hides failures. Set targets per slice: clean printed text, handwriting, digits and Latin runs, phone photos. Compare the vendor's ground truth with your own double transcription on the same pages; disagreements on digits, dots and reading order matter more than a small overall CER.
Can I use public Arabic OCR datasets commercially?
Only if each dataset's own license allows it. Many research datasets restrict commercial use or redistribution. Muharaf, for example, distributes its public part under CC BY-NC-SA 4.0 and a restricted part under a proprietary license limited to research use, with no redistribution. Read the license file for every source, and keep training, display and redistribution rights separate.
Should synthetic rendered pages count toward dataset size?
Report them separately. Rendered pages are useful for fonts and rare characters, but they do not reproduce real scanners, stamps, folds or handwriting. Ask for genuine, de-identified and synthetic pages to be labeled as such, and evaluate only on genuine pages unless your production traffic is synthetic too.
What metric should I use for invoice and form extraction?
Score each field, not the page. Use exact match on normalized values (dates, amounts, ID numbers) and report precision, recall and F1 per field type, plus key-value linking accuracy. Keep a separate CER on the raw text, so you can tell an OCR miss from an extraction miss.