Solutions · العربية
Arabic OCR and manuscript reading
Printed Arabic measured honestly on the free tier — and a measured 95.3% on a real handwritten manuscript where classical OCR collapses.
Arabic OCR quality claims are rarely accompanied by measurements, and the hard cases — Arabic-Indic numerals on invoices, diacritised text, and above all handwriting — are exactly where tools fail silently. For archives and heritage collections, "supported" usually means "produces something".
We publish the real numbers for both tiers, including the case that matters to archives: a genuine 19th-century handwritten manuscript page, where the difference between engines is not quality but usability.
Print and manuscript, both measured
The free figure, the failure case, and the AI-tier result on real handwritten material.
Measured 2026-08-03
- Printed Arabic, free tier
- 93.84% character accuracy on modern print — measured on a paragraph deliberately containing Arabic-Indic numerals (٢٤١٧, ٢٠٢٦), the digits invoices carry and the likeliest characters to misread. Errors cluster in dots and digit shapes, not word structure.
- A real handwritten manuscript
- A 19th-century naskh manuscript page: classical OCR read it at 43.9% — honestly unusable, which is what every conventional tool will give you. The AI tier read the same page at 95.3%, opening lines letter-perfect, at roughly a tenth of a cent per page.
- What that enables
- Archive and heritage digitisation at a cost that changes the project maths: a 10,000-page manuscript collection reads for around ten dollars of model cost, with output a scholar corrects rather than retypes.
- Right-to-left, done properly
- Output returns in UTF-8 logical order — it pastes correctly into Word, Docs and any RTL-aware system, with Latin fragments and digits inside Arabic text handled in one pass.
- Where your documents live
- Processed on our own servers in Frankfurt, Germany. Originals are deleted within an hour of processing, results within 30 days, and nothing is ever used to train models — ours or anyone else’s. The security page and DPA spell out the rest.
Method and raw results: how we measure. Every figure is dated and reproducible.
How you use it
- Print, free today
- The Arabic image-to-text tool runs genuine Arabic recognition free, no account, 5 files a day.
- Invoices and business documents
- Structured extraction reads Arabic-content documents into fields with arithmetic checks — and flags the Arabic-Indic digits it is unsure of instead of guessing silently.
- Manuscripts and archives
- Try a page yourself on the handwriting-to-text tool — the measured 95.3% capability, one credit a page. For a collection it becomes a conversation: send three representative pages, we measure them, and the quote rests on numbers we have both seen.
What to know before you commit
The limits, stated here rather than discovered in your evaluation.
Fully diacritised text scores below the print figure. Harakat are the smallest marks on the page and vanish first with blur or low resolution. Qurʼanic and poetry material deserves a measured sample, not an assumption.
The manuscript figure is one page, not a corpus. One real manuscript measured at 95.3% is evidence, not a statistic — hands, inks and centuries vary. For a collection, we measure your material first; that is the difference between a claim and a quote.
Questions buyers ask
- Can it really read handwritten manuscripts?
- On the page we measured — genuine 19th-century naskh — yes: 95.3% on the AI tier, with the opening lines letter-perfect. Classical OCR managed 43.9% on the same page. One page is evidence not proof, so manuscript projects start with a free measurement of your material.
- Are Arabic-Indic numerals (٠١٢٣) handled?
- Yes, and honestly: they are inside our measured 93.84% print figure and they are the likeliest characters to misread. On financial documents the extraction tier cross-checks digits against the document’s own arithmetic.
- Does the output paste correctly right-to-left?
- Yes — UTF-8 logical order, which every RTL-aware application renders correctly. No reversed strings, no manual reordering.
- We are a library or archive — how do we start?
- Email three representative pages. We measure them through both tiers, send you the numbers with the method, and quote from there. Processing is in Frankfurt with a published deletion schedule and no training on your material.
Prefer a conversation first? Write to the people who built it
Related: Handwriting transcription · Free Arabic tool · Invoice OCR · How we measure