Invoices and receipts
Arabic-English line items, totals, taxes, discounts, and stamps, with the scan quality and layout variation your system encounters.
Build a dataset for the documents your product reads, from bilingual invoices to handwritten forms. Get text, layout, extracted fields, and evidence for the answers you need.
Arabic-English line items, totals, taxes, discounts, and stamps, with the scan quality and layout variation your system encounters.
Printed fields, handwriting, checkboxes, corrections, and missing entries in applications, delivery documents, or service reports.
Clauses, tables, attachments, and cross-references, with questions answered from specific supporting passages.
Choose reading order, layout, normalized fields, evidence locations, and checks for missing information or inconsistent totals.
We define source permissions, layout diversity, capture conditions, and page count around the problem you need to solve.
Genuine, de-identified, and synthetic documents remain distinguishable. Template diversity is reported separately from page count.
Evaluation splits can separate issuers, templates, sources, and writers to check performance on documents the model has not seen.
A bilingual invoice has a discount and a handwritten correction. Labels locate both amounts so the system can extract them, recompute the total, and flag a mismatch.
Yes. You can request transcription, coordinates, and reading order on their own, or add layout, field extraction, document questions, and validation checks. The labels should match the model you are building.
For a custom collection, we agree the template coverage and evaluation splits with you. Variations of one template remain identifiable, so page count does not get mistaken for the number of independent layouts.