Solutions · සිංහල
Sinhala OCR
The language most OCR vendors skip, measured honestly — free tier and AI tier, on real Sinhala paper.
Sinhala is spoken by twenty million people and ignored by nearly every OCR vendor: not offered, or listed but never measured. Sri Lankan businesses, archives and government offices digitise documents by typing them.
VisionParse is built by a Sri Lankan company and treats Sinhala as a first-class language: a free tool anyone can use today, an AI tier for serious material, and published measurements for both — including on a real, degraded book scan, not just a clean render.
Two tiers, both measured on Sinhala
The free figure and the paid figure, with the difference stated in numbers rather than adjectives.
Measured 2026-08-03
- Free tool, clean print
- 97.48% character accuracy on printed Sinhala, including combining vowel signs (පිලි) and a conjunct with a zero-width joiner — the character class that breaks naive pipelines. Free, no account, 5 files a day.
- Free tool, a real degraded book scan
- 94.5% on a genuinely degraded book page a customer supplied. Every error was a similar-loop letter pair (ච/ව, න/ත, ස/හ) — exactly the failure mode we document, and the confidence score flagged it honestly at 91.9%.
- AI tier, the same real scan
- 99.16% on the identical page through the model our paid extraction uses — it repaired essentially every loop confusion, in 2 seconds, for a fraction of a cent. For archival and degraded material, the gap between tiers is measured, not claimed.
- Mixed-script reality
- Sri Lankan documents mix Sinhala, English and digits constantly; the recognition reads all three in one pass, and output is standards-compliant Unicode that pastes anywhere.
- Where your documents live
- Processed on our own servers in Frankfurt, Germany. Originals are deleted within an hour of processing, results within 30 days, and nothing is ever used to train models — ours or anyone else’s. The security page and DPA spell out the rest.
Method and raw results: how we measure. Every figure is dated and reproducible.
How you use it
- Try it free right now
- The Sinhala image-to-text tool runs genuine Sinhala recognition — not English guessing at the script — with no account, on 5 files a day.
- For serious volumes
- Paid plans meter by page with the API included; the AI tier handles the degraded material the free engine loses. Books, archives, ledgers — send us a sample and we will measure it before you commit.
- For institutions
- Documents are processed in Frankfurt, deleted on a published schedule, and never used for training — the answers a library or ministry asks first, written down before they ask.
What to know before you commit
The limits, stated here rather than discovered in your evaluation.
Low-resolution newsprint is the free tier’s hard case. Sinhala letters differ by small loops, and below roughly 30 pixels of letter height those features merge. The AI tier recovers most of it — that is precisely the measured 94.5% → 99.2% gap.
Handwritten Sinhala is unmeasured. We have not benchmarked Sinhala handwriting yet, so we make no claim for it. Our Arabic manuscript result suggests the AI tier is the place to try; send a sample and we will measure it rather than guess.
Questions buyers ask
- Why should we trust the accuracy claims?
- Because they are dated measurements with published method and known limits — including the numbers that are not perfect. The 94.5% on a real degraded scan is on this page next to the 99.2%; a vendor who only shows you round numbers is showing you marketing.
- Can it digitise old Sinhala books?
- This is exactly the measured case: a real degraded book page read at 94.5% free and 99.2% on the AI tier. For a digitisation project, send three representative pages and we will benchmark them before you spend anything.
- Does copy-paste work properly in Word and Docs?
- Yes — output is standards-compliant Unicode Sinhala, conjuncts included, verified at the codepoint level. If it renders in your editor, it pastes correctly.
- Is there an API for Sinhala?
- Yes, the same API as everything else: pass sin as the language for OCR, and the extraction endpoint handles Sinhala content inside structured documents.
Prefer a conversation first? Write to the people who built it
Related: Free Sinhala tool · Tamil OCR · The OCR API · How we measure