Solutions · தமிழ்
Tamil OCR
Eighty million speakers, and OCR vendors treat the script as an afterthought. We measure it and publish the number.
Tamil is one of the world’s major languages, used across Tamil Nadu, Sri Lanka, Singapore and a global diaspora — and mainstream OCR either omits it or lists it unmeasured. Offices process Tamil paperwork by retyping it.
We measured our Tamil recognition and published the figure with its limits: 95.35% character accuracy on print, with the two-sided vowel signs and ligatures that make the script genuinely hard handled correctly.
The measured state of Tamil support
What we measured, what came back correct, and what we have not measured yet.
Measured 2026-08-03
- Print accuracy
- 95.35% character accuracy on printed Tamil, scored character-by-character against exact ground truth. Errors concentrate in similar curved letter pairs (ன/ண, ர/ா) at small sizes — sharper sources close the gap.
- The script’s hard parts held
- Vowel signs attaching before, after and around consonants — including forms that fully wrap their consonant — stayed attached to the right letters and returned as correct Unicode combinations, so the text is searchable, not just readable.
- Mixed-script documents
- Tamil invoices and forms mix Tamil, English and digits; recognition reads all three in one pass.
- The AI tier, honestly
- For Sinhala and Arabic we have measured the AI tier rescuing degraded material (94.5%→99.2%, and 43.9%→95.3% on a manuscript). We have not yet run that comparison for Tamil — send us a real degraded Tamil scan and we will measure it rather than assume the pattern transfers.
- Where your documents live
- Processed on our own servers in Frankfurt, Germany. Originals are deleted within an hour of processing, results within 30 days, and nothing is ever used to train models — ours or anyone else’s. The security page and DPA spell out the rest.
Method and raw results: how we measure. Every figure is dated and reproducible.
How you use it
- Try it free right now
- The Tamil image-to-text tool runs genuine Tamil recognition, free, no account, 5 files a day.
- For volume
- Paid plans meter by page with API access; pass tam as the language code. Archives and ledgers deserve a measured sample first — send three pages.
- For institutions
- Frankfurt processing, published deletion schedule, no training on customer documents — the compliance answers, pre-written.
What to know before you commit
The limits, stated here rather than discovered in your evaluation.
Grantha letters and dense older print score below the figure. The Sanskrit-loan letters (ஜ, ஷ, ஸ, ஹ) and tightly-set older books are harder than modern print. If your material is archival, benchmark a sample before committing a workflow.
No degraded-scan measurement for Tamil yet. We say this plainly rather than borrowing the Sinhala result: the AI-tier rescue is measured for Sinhala and Arabic, unmeasured for Tamil. The first real Tamil scan a customer sends us gets measured and published here.
Questions buyers ask
- Why is the published figure 95.35% and not 99%?
- Because it is measured. Most errors are similar-shaped letter pairs at small sizes, and sharper images push accuracy up — but we publish what we scored, not what we hope. The known-limits list says exactly where it degrades.
- Does it handle Sri Lankan Tamil documents?
- Yes — same script, and mixed Tamil-Sinhala-English paperwork is normal here. We are a Sri Lankan company; this is home-market material for us.
- Can it read Tamil handwriting?
- Unmeasured, so unclaimed. Our Arabic manuscript result (95.3% on the AI tier) suggests where to try, and we would measure your sample before you spend anything.
- Is Tamil available over the API?
- Yes — pass tam as the language code on the OCR endpoint. Same keys, same limits, same measured honesty.
Prefer a conversation first? Write to the people who built it
Related: Free Tamil tool · Sinhala OCR · The OCR API · How we measure