Receipt Datasets for Expense and Retail AI

Annotated receipt corpora across retail formats, geographies, and capture conditions with merchant, item, and total field labels. Built for expense management AI and retail analytics.

Abstract data visualization representing receipt scanning
20+
Years in AI training data
1,000+
Language locales
1M+
Hours of video annotated
ISO
27001 certified

Beyond Public Receipt Benchmarks

CORD and SROIE established receipt parsing baselines. But production expense AI and retail analytics systems handle receipt diversity, low-quality mobile captures, and multilingual formats those benchmarks do not represent.

Expense management AI must parse receipts from restaurants, hotels, travel, fuel, and professional services across global markets. Retail analytics models must identify item-level detail, discounts, and loyalty point transactions. Generic benchmarks cover none of this depth.

Off-the-shelf receipt datasets suffer from capture quality gaps (benchmark datasets use controlled scans; real users photograph crumpled receipts in poor light) and format diversity gaps across retail verticals.

LXT builds custom receipt datasets matched to your receipt population. We annotate merchant information, line items, payment methods, and tax fields across the receipt types, capture conditions, and languages your system processes.

Limitations of Public Receipt AI Datasets

Standard benchmarks serve research well. Production deployments need more.

DatasetPrimary LimitationImpact
CORDKorean restaurant receipts only; narrow geographic and vertical coverage; limited to simple single-block layouts without complex itemizationSingle vertical
SROIE419 English receipts only; simple four-field extraction task; insufficient scale and diversity for production expense AILimited scope
DocVQAMixed document types with VQA framing; not optimized for receipt-specific field extraction; limited receipt examplesVQA framing
XFUNDForm understanding focus; non-receipt document types; annotation schema mismatches expense extraction needsWrong schema
EATENEntity-aware scene text; mixed receipt and non-receipt content; limited annotation depth for financial fieldsMixed content

Not sure which specs you need?

Our data specialists help you scope the right dataset for your model architecture.

Talk to a Specialist

Specs Built Around Your Model

Public datasets come fixed. Yours is configured for your architecture, environment, and use case.

Receipt Types

Vertical Coverage

  • Retail Verticals: Grocery, restaurant, fuel, pharmacy, hotel, and travel receipts
  • Payment Types: Cash, card, contactless, mobile payment, and split-payment receipts
  • Languages: Multilingual receipts across 30+ countries and tax regimes

Capture Conditions

Image Quality

  • Mobile Photography: Handheld smartphone captures with natural lighting variation
  • Scan Quality: Office scanner, receipt scanner, and photographed receipt variation
  • Physical Condition: Crumpled, faded thermal paper, torn, and ink-faded receipts

Annotation Fields

Extraction Targets

  • Merchant Fields: Business name, address, phone, tax ID, and receipt number
  • Line Items: Item name, quantity, unit price, and line total per purchased item
  • Payment Summary: Subtotal, tax, tip, discount, total, and payment method fields

Need a custom configuration?

We've built datasets across dozens of domains and use cases. Let's scope yours.

Get a Custom Quote

Edge Cases in Receipt Processing

High-accuracy models handle rare attributes that public datasets miss.

Thermal Paper Degradation

Thermal receipt paper fades rapidly. Models must handle near-invisible text on receipts that have been exposed to heat or light, a common real-world capture scenario.

Multiple Receipts in One Image

Expense app users sometimes photograph multiple receipts together. Receipt boundary detection and individual receipt segmentation support multi-receipt capture workflows.

Handwritten and Mixed Print Receipts

Small vendors issue handwritten receipts or receipts mixing printed headers with handwritten amounts. Mixed print and handwriting annotation supports universal receipt parsing.

International Tax and Fee Structures

Service charges, VAT, GST, and local taxes appear in different positions and formats by country. Jurisdiction-specific field annotation ensures correct tax extraction globally.

Human-in-the-Loop Annotation

Precise annotation bridges raw data and learnable signal. Expert annotators deliver precision automated tools can't match.

🧾

Receipt Field Annotation

Expert annotators mark all merchant, item, and payment fields with bounding boxes and normalized values. Multi-pass verification ensures extraction accuracy.

💸

Line-Item Extraction

Item-level annotations with description, quantity, and price fields supporting both structured (barcode-style) and unstructured (handwritten) receipt formats.

📸

Image Quality Metadata

Blur score, brightness, skew angle, and paper degradation level annotated per receipt for quality-aware model training and capture feedback systems.

Receipt Datasets for Your Domain

Custom taxonomies and collection protocols for specific deployment contexts.

💼

Expense Management

Corporate expense capture, reimbursement automation

🛒

Retail Analytics

Item-level purchase analysis, basket intelligence

💳

Loyalty Programs

Purchase verification, points calculation, redemption

🧾

Fraud Detection

Receipt forgery detection, duplicate submission flags

📱

Mobile Finance Apps

Budgeting apps, spending categorization, receipt storage

🏦

SME Accounting

Automated bookkeeping, VAT reclaim, tax filing

🛍️

E-Commerce Returns

Return verification, refund processing, warranty claims

🌍

Global Travel

Multi-currency expense parsing for business travel

Secure and Ethical Data Collection

Data collection involving people and sensitive content requires robust security, compliance, and ethical protocols at every stage.

🌐

Global Demographic Reach

Collection across 1,000+ locales and diverse demographics to prevent algorithmic bias in your deployed models.

🔒

ISO 27001 Certified

Sensitive projects processed in certified secure facilities meeting the highest information security standards.

✅

GDPR & Privacy Compliance

All collection and annotation protocols vetted for consent and privacy. Legally robust for global deployment.

Receipt Dataset FAQs

What receipt types and retail verticals can you cover?+
We cover grocery, restaurant, fuel, pharmacy, hotel, travel, and general retail receipts. Custom verticals are supported with targeted collection from relevant merchant categories.
Can you handle poor-quality smartphone photographs?+
Yes. We specifically design collection protocols to include real smartphone captures under variable lighting, angles, and physical receipt conditions matching your user base.
Do you support multilingual receipt annotation?+
Yes. We annotate receipts in 30+ languages with tax field schemas adapted to local regulations. Language-specific annotators handle non-Latin scripts including Arabic, Chinese, and Japanese.
What output format does the annotation use?+
CORD-compatible JSON, COCO format, and custom key-value JSON. Normalized field values include ISO date formats, float amounts, and merchant name standardization where requested.
Can you annotate receipts from our existing app captures?+
Yes. We can annotate receipts submitted through your app or platform under a data processing agreement. We handle anonymization of cardholder information before annotation begins.
What does a custom receipt dataset cost?+
Projects range from $8K for focused single-vertical datasets (1,000-5,000 receipts) to $60K+ for large multi-vertical, multi-language corpora.
How do you handle cardholder and PII data on receipts?+
Card numbers, names, and other PII visible on receipts are redacted or pseudonymized before or during annotation depending on your requirements. All handling follows GDPR and applicable privacy standards.

Scope Your Custom Receipt Dataset

Share your receipt types, capture conditions, and extraction targets. A document AI specialist will provide a detailed annotation plan within 48 hours.

Contact us.

Please provide us with the details of your inquiry and one of our team members will be in touch.

Join our global team of contributors today

Apply here to be considered for future projects including data collection, annotation and transcription
Start application
(opens in a new tab)