Document problems, measured answers
Each page below carries its own dated benchmark figures — including the imperfect ones, because that is what makes the good ones believable.
Free tools for one-off conversions live on the tools hub. These pages are for teams putting documents through every day.
The split is about what you are deciding. A tool page exists to get one file converted in the next minute — free, no account, and you leave with the text. A solution page exists because someone is choosing what their company will run invoices, receipts or Arabic paperwork through for the next two years, and that decision needs different evidence: what it costs at volume, what it does when a document is unusual, and where it has been measured to fail.
So each page below leads with a dated figure from our benchmark log, names the limit beside it, and links to the free tool that does the same job so you can test the claim before you talk to anyone. Where a capability runs on the paid AI engine it says so, and what a document costs is on the pricing page rather than behind a form.
- Invoice OCR and data extractionSupplier invoices become checked fields — 76/76 benchmark fields exact, real purchase orders read to the cent.
- Receipt OCR and expense extractionThermal paper to expense data — 14/14 fields exact, and exact again faded, creased and photographed at an angle.
- The OCR APITwo calls, no SDK — 0.5–0.65 s per page measured, keys stored only as hashes, honest machine-readable limits.
- Sinhala OCRThe language most vendors skip — 97.5% on print free, a real degraded book scan at 99.2% on the AI tier.
- Tamil OCR95.35% measured on print, ligatures handled — and the gaps stated instead of glossed.
- Arabic OCR and manuscript readingPrint at 93.8% free — and a real 19th-century manuscript at 95.3% on the AI tier, where classical OCR managed 43.9%.