Last updated
Document problems, measured answers
Each page below carries its own dated benchmark figures — including the imperfect ones, because that is what makes the good ones believable.
Free tools for one-off conversions live on the tools hub. These pages are for teams putting documents through every day.
- Invoice OCR and data extractionSupplier invoices become checked fields — 76/76 benchmark fields exact, real purchase orders read to the cent.
- Receipt OCR and expense extractionThermal paper to expense data — 9/9 fields exact on our thermal-layout benchmark, items included.
- The OCR APITwo calls, no SDK — 0.5–0.65 s per page measured, keys stored only as hashes, honest machine-readable limits.
- Sinhala OCRThe language most vendors skip — 97.5% on print free, a real degraded book scan at 99.2% on the AI tier.
- Tamil OCR95.35% measured on print, ligatures handled — and the gaps stated instead of glossed.
- Arabic OCR and manuscript readingPrint at 93.8% free — and a real 19th-century manuscript at 95.3% on the AI tier, where classical OCR managed 43.9%.