How to audit an Arabic image and video dataset for a multimodal model: 10 checks
Draw your own sample, open the images and clips, and score them check by check. This audit covers regional representation, whether the Arabic in each image can actually be read, label agreement, task fit, collection consent and leakage into the test set, with the public resources to compare against.
To audit an Arabic image or video dataset for a multimodal model, draw your own sample of a few hundred images or clips, have native readers re-label part of it, and score ten things separately: task fit, regional representation, capture conditions, whether the Arabic text in each image is legible, transcription rules, grounding evidence, label agreement, video timing, consent and rights, and leakage between training and test data. Never accept one blended accuracy number. Arabic visual data fails in specific places: a sign transcribed in reverse, an answer the image does not support, the same storefront in both the training and the test split.
This checklist is for vision-language and multimodal teams in the UAE, Saudi Arabia and elsewhere that are buying, licensing or building on Arabic image and video data: street and retail scenes, packaging, menus, signs, instructions and authorized video. For scanned documents, invoices and forms, use our Arabic OCR dataset audit instead; documents need their own checks for reading order, fields and tables.
- Start from the task. Captions, scene-text reading, grounded question answering and video events need different labels. A dataset built for one rarely serves another.
- Ask where, not just how many. Count countries, cities, settings and capture devices. A set collected somewhere else will not show your users' signs, shelves and packaging.
- Every answer needs visible evidence. If a question can be answered without looking at the image, it tests the language model, not vision.
- Consent and splits decide whether you can ship. People, faces and number plates are personal data, and near-duplicate scenes across splits inflate every score.
Why Arabic visual data needs its own audit
Public benchmarks show that current models still struggle with Arabic visual content, so the data used to train and test them has to be right where they fail. CAMEL-Bench (Ghaboura et al., 2024), an Arabic benchmark of around 29,036 questions across eight domains and 38 sub-domains, including multi-image and video understanding, reports that even GPT-4o reached an overall score of 62%. JEEM (Kadaoui et al., 2025), which tests image captioning and visual question answering in the dialects of Jordan, the Emirates, Egypt and Morocco, found that the five open-source Arabic vision-language models it evaluated consistently underperformed, struggling with both visual understanding and dialect-specific generation, and that even the best model, GPT-4V, varied in linguistic competence across dialects and lagged in visual understanding.
Three facts drive most of the checks below:
- Where an image was taken changes what a model sees. DeVries et al. (2019) found that commercial object-recognition systems performed relatively poorly on household items common in countries with low household income, with qualitative analysis suggesting this was mainly because of differences in appearance within an object class and items appearing in a different context.
- Arabic text in scenes is often mixed and hard to read. Signs, packaging and menus mix Arabic, Latin script and digits in the same frame, under glare, angle and distance. Unicode stores Arabic in reading order and the Bidirectional Algorithm reorders it for display, so text exported in visual (display) order can look right in a tool that does not apply the algorithm and still be stored backwards.
- Test sets leak. Barz and Denzler found that 3.3% of the CIFAR-10 and 10% of the CIFAR-100 test images have near-duplicates in the training set, and Song et al. (Findings of EMNLP 2025) report significant contamination across twelve multimodal LLMs and five benchmarks, particularly in proprietary models and older benchmarks.
The 10-point Arabic image and video dataset audit
Run every check on the same sample, drawn by you rather than the vendor and stratified by scene type, location and capture condition. A red flag on checks 1, 2, 6, 9 or 10 usually ends the evaluation; the others can often be fixed with a specification change and a re-delivery.
| # | Check | How to inspect | Red flag |
|---|---|---|---|
| 1 | Task fit | Map each label type to the behavior you need in production: read text, answer a question, point to evidence, describe a scene, follow an event in video. | Captions only, when you need answers grounded in a region; English questions machine-translated into Arabic. |
| 2 | Regional representation | Ask for per-image capture metadata: country, city, setting (street, mall, pharmacy, supermarket, transport) and source. Compare it with where your users are. | Only a total image count; web-scraped stock photos; one city standing in for "the Arab world". |
| 3 | Capture conditions | Check the spread of lighting, distance, angle, motion blur, glare on glass and packaging, night scenes and device types. | Only clean, well-lit, front-on photos when your users point phones at shelves and signs. |
| 4 | Arabic text legibility | For every text region, ask for a legibility flag (legible, partly legible, not legible) and a script tag (Arabic, Latin, mixed). Crop 50 regions and read them yourself. | Full transcriptions attached to text no reader could make out; no way to exclude unreadable regions from scoring. |
| 5 | Transcription rules | Check that text is stored in logical (reading) order, digits as printed, and letters as standard Unicode rather than presentation forms (the same checks as in our OCR audit). | Reversed strings such as ةيلديص for صيدلية ("pharmacy"); Arabic-Indic digits silently converted. |
| 6 | Grounding evidence | For question-answer pairs and descriptions, check that each answer points to a box, region or timestamp, and that the cited region actually contains the answer. | Answers that a text-only model could guess; evidence boxes on the wrong object; answers that rely on knowledge outside the image. |
| 7 | Label consistency | Ask for double-annotation and adjudication records. Re-label 30 to 50 items yourself: transcriptions, answers, boxes. | No agreement figures; your re-labels disagree on what the text says or where the object is. |
| 8 | Video timing and scope | For clips, check event start and end times, the tolerance used, and whether questions about order or change match the footage. Confirm whether audio or transcripts are in scope. | One label per clip when you need events; timestamps that drift; speech questions in a "vision" set with no transcript. |
| 9 | Consent and rights | Ask how each image and clip was obtained, what permissions cover people, faces, number plates and private premises, and what the license allows: training, display, redistribution. | No written source permissions; identifiable people with no consent record; attributes such as nationality or religion inferred from appearance. |
| 10 | Test leakage | Search for near-duplicates across splits (the same storefront from two angles, neighboring video frames) and check that public benchmark images are not inside the training set. | Random splits over frames or photo bursts; no location- or source-disjoint test set; no duplicate check. |
Worked example: one pharmacy photo, three labels that fail
Take one phone photo of a bilingual pharmacy storefront at night (an illustrative case, not a delivered record). The sign reads صيدلية above "PHARMACY", and a smaller board beside the door lists opening hours, half hidden by glare.
| Delivered label | What is wrong | Which check fails |
|---|---|---|
| Transcription ةيلديص | Stored in display order, so search and scoring see a different word | 5: transcription rules |
| Q: "When does it close?" A: "Open 24 hours", no evidence box | The hours board is not legible in this photo; the answer came from the annotator, not the image | 4 and 6: legibility and grounding |
| The same storefront, photographed a minute later from another angle, in the test split | The model has seen this sign in training; the test score overstates real-world reading | 10: test leakage |
None of these errors shows in an overall accuracy figure. All three show up within minutes when you render the transcription right to left beside the crop, ask for the evidence box, and check where the photos were taken.
Which metric for which task
Ask the vendor which of these they measured on their own labels, on what sample, against whose second annotation. Then compute the same numbers on your sample, by slice: country or city, setting, capture condition, Arabic versus mixed-script text.
| Task | Metric | Arabic-specific detail |
|---|---|---|
| Scene-text detection | Precision, recall and their harmonic mean (H-mean) over text regions matched by overlap (commonly IoU of 0.5 or more) | Score Arabic, Latin and mixed regions separately; exclude regions flagged not legible instead of counting them as misses. |
| Scene-text recognition | Character error rate and word accuracy on cropped regions | Report with and without diacritics, and separately for digits; normalize letter variants (alef forms, taa marbuta) only at scoring time, never in the stored label. |
| Visual question answering | VQA accuracy, from Antol et al. (ICCV 2015): 10 human answers per question, and an answer counts as fully correct if at least 3 people gave it | Arabic answers vary in form (with or without the article ال, numerals in either digit set), so collect several reference answers; use exact match for prices, times and quantities. |
| Grounding | Share of answers with an evidence region; overlap (IoU) between the cited region and a reference region | Check that the cited region holds the Arabic text the answer quotes, not only the object. |
| Video events | Temporal overlap between labeled and reference event segments; accuracy on order and change questions | State the boundary tolerance; keep speech-based questions in a separate, transcript-backed slice. |
| Label agreement | Chance-corrected agreement (kappa or alpha) on categorical labels; box overlap and transcription CER between two annotators | Artstein and Poesio (2008) survey which coefficient to use when. |
Public resources to compare against
Public sets are a useful second opinion: run the same checks on a public sample and on the offered dataset, and compare. Read each license before any commercial use; many research sets do not allow it.
| Resource | What it covers | Use in an audit |
|---|---|---|
| CAMEL-Bench (2024) | Around 29,036 Arabic questions, 8 domains and 38 sub-domains, including video understanding, handwritten documents and remote sensing; manually verified by native speakers | A task list to map your checks against; a reference for where models fail |
| JEEM (2025) | Captioning and visual question answering in four dialects: Jordan, the Emirates, Egypt and Morocco | A public reference that includes Emirati-dialect captions and answers |
| PEARL (2025) | Over 309,000 culturally focused multimodal examples across ten domains covering all Arab countries, with the PEARL, PEARL-LITE and PEARL-X benchmarks | Cultural-fit questions to compare with; the dataset card lists CC BY-NC-ND 4.0, so no commercial use or derivatives |
| Henna (Peacock, ACL 2024) | A benchmark for multimodal models on Arabic culture; per the project README, images from 11 Arab countries with QA pairs generated by GPT-4V from each image and its Wikipedia article | A cultural test; note that its answers are model-generated, which is itself a check-7 question |
| EvArEST (IEEE Access 2021) | 510 Arabic-English scene images with word polygons and language tags, 7,232 cropped word images, and about 200,000 synthetic images; BSD-3-Clause | Scene-text labeling conventions and a small test of reading in real scenes |
| ICDAR 2019 MLT (Nayef et al., 2019) | 20,000 real images with text in 10 languages, Arabic among them, for detection, script identification and end-to-end reading | Detection and end-to-end protocols; mixed-script scenes |
Consent and personal data in UAE and Saudi collections
Street, retail and video data almost always contains people, faces, number plates or private premises. In the UAE, Federal Decree-Law No. 45 of 2021 on personal data protection, in force since 2 January 2022 according to the UAE government portal, defines biometric data as personal data produced by technical processing of physical characteristics, with facial images as an example, and lists biometric data among sensitive personal data. Article 4 prohibits processing personal data without the owner's consent, subject to listed exceptions. Free-zone companies with their own data protection legislation are excluded (Article 2): the same portal cites the DIFC's Data Protection Law No. 5 of 2020, and the ADGM applies its Data Protection Regulations 2021. In Saudi Arabia, the Personal Data Protection Law and its implementing regulation, listed by SDAIA, set the rights of data subjects and the obligations of controllers.
For an audit this means three questions, answered in writing before you license anything: what consent or other lawful basis covers each identifiable person, whether faces and plates were blurred or kept (and why), and which law applies where the data was collected and where it will be stored. This is not legal advice; confirm the position with counsel for your entity and free zone.
Ask for the dataset's documentation
Datasheets for Datasets (Gebru et al.) proposes documenting each dataset's motivation, composition, collection process and recommended uses. For an Arabic visual dataset, the datasheet should answer at least:
- The task each label type was built for, and the label schema with its version.
- Counts by country, city, setting, source and capture device, separately from the image or clip count.
- Which items are field-captured, licensed from third parties, scraped or generated, labeled per item.
- The transcription and legibility rules, and how unreadable text is excluded from scoring.
- Who annotated and who adjudicated, the share double-annotated, and the agreement on that share.
- Split rules, the near-duplicate check, and whether test locations or sources are kept out of training.
- Consent records, how faces, plates and other personal data are handled, and the license, with training, display and redistribution stated separately.
For residency and access questions when the data includes people, see our PDPL and data residency checklist. For licensing of off-the-shelf data more broadly, see commercial Arabic datasets.
When no offered dataset passes
Often none will, because the scenes you need are specific: one retailer's shelves, one city's road signs, one product line's packaging, one operator's authorized camera footage. Generated images can add variety in fonts and layouts, but they do not reproduce real glare, distance or wear, and they should never be counted as field collection (see the limits of synthetic data). The practical answer is a custom collection or annotation built against this checklist from the first image. The vendor questions that matter are in how to choose an Arabic data annotation vendor.
Why vision teams build Arabic image and video data with Bayanat Labs
When the visual data you need does not exist off the shelf, Bayanat Labs is the best partner to build it, because the hardest points in this audit, from grounding and capture planning to consent and evaluation splits, are part of how we deliver:
- Built around your task. Bayanat Vision is a custom scope: we start with one buyer problem, agree source permissions and a collection plan for a small pilot, then set the image or video count and coverage with you.
- Answers tied to visible evidence. Specifications can include scene text, coordinates, object relationships and answers linked to visible evidence, and video projects can add event labels and timestamps (checks 4 to 8).
- Capture planned, not left to chance. Lighting, capture conditions, text styles, layouts and difficulty are planned around the task, and generated images stay identifiable separately from field collection (checks 2 and 3).
- Consent and splits handled up front. Source and media permissions are agreed before collection, we describe the collection without inferring people's identities from appearance, and evaluation can use separate locations or sources (checks 9 and 10).
- Native readers with quality control. Vetted native speakers and linguists across 25+ Arabic varieties, screened for dialect and domain before they touch client data, calibrated against gold standards with adjudication and rework loops, and an audit trail on every task (how our annotation runs).
- Sovereign by default, proof before scale. In-region hosting by default with on-prem and private-cloud options, and a fixed-scope pilot with gold-standard QA and a quality report, usually within two weeks.
We apply the same discipline to automated labeling: our LabelBench audit shows how a high binary accuracy can hide failure on the real task, which is why we report by slice.
Frequently asked questions
How many images do I need to audit an Arabic image dataset?
Enough to see every slice you care about. A practical start is a few hundred images or clips that you draw yourself, not the vendor, stratified by country or city, setting and capture condition, with re-labeling of 30 to 50 items by your own native readers to measure label quality, not just model quality.
How do I check that an Arabic VQA answer is grounded in the image?
Ask for an evidence region or timestamp for every answer, then open a sample and confirm the region contains the text or object the answer relies on. Also try the questions without the image: if a text-only model answers them well, the set is testing language priors, not vision.
Can I use public Arabic multimodal datasets commercially?
Only if each dataset's own license allows it. PEARL, for example, is listed under CC BY-NC-ND 4.0, which excludes commercial use and derivatives, while EvArEST is under BSD-3-Clause. Read the license for every source, and keep training, display and redistribution rights separate.
Should generated images count toward dataset size?
Report them separately. Generated images can add fonts, layouts and rare scenes, but they do not reproduce real glare, distance, wear or camera noise. Ask for field-captured, licensed and generated items to be labeled per item, and evaluate on field-captured items unless your production traffic is generated too.
How should I split an Arabic image or video dataset to avoid leakage?
Split by location, source or recording, not by individual image or frame. Photo bursts, the same storefront from two angles and neighboring video frames must stay on one side. Run a near-duplicate check across splits, and check that public benchmark images are not in your training data.