Skip to content
DocuExtract

Languages

35+ scripts. Real coverage.
Best-in-class engine per language.

We route each script to the engine that actually handles it well — PaddleOCR for printed CJK, Surya for Indic + Arabic, Tesseract for Latin, vision-LLM for handwriting. Three maturity tiers (stable / beta / experimental), per-language accuracy notes, honest caveats. The full matrix below.

Beyond language

We also read what others miss.

Language coverage is only half the story. The other half is document shape and quality: handwriting, complex tables, degraded scans. Here's what the engine handles regardless of language.

Printed text & born-digital PDFs

Embedded PDF text (born-digital) reads at 100% accuracy with no model involved. Printed scans run through Tesseract for Latin scripts and PaddleOCR for CJK/Indic.

Handwriting — three tiers

Handwritten English (Latin/Cyrillic/Greek) uses a fast vision model at 5 cr/page. Handwritten non-English (CJK/Indic/Arabic) routes to the script's best-in-class model — Gemini Flash for CJK, GPT-4o for Arabic/Hebrew. Handwritten mixed-language (forms with 2+ scripts) runs a dual-engine pipeline. Each tier priced separately.

Multi-column tables

Single-page tables extracted with bounding-polygon awareness. Multi-page tables (continuation rows across pages) are on the roadmap — tracked in KNOWN_ISSUES.

Degraded scans

Low-resolution photos, faxed documents, stained pages. The cascade escalates to vision-LLM automatically when traditional OCR confidence drops below 0.65.

Tier · Stable

Stable support

Production-ready. Validated against a comprehensive fixture set. Use these in business-critical workflows with confidence.

English

en
Script
Latin
Engine
Tesseract
Accuracy
99%+ on born-digital PDFs; 96–98% on clean scans.

Caveats

Degraded scans escalate to vision-LLM (Tier 3).

Spanish

es
Script
Latin
Engine
Tesseract
Accuracy
99%+ on born-digital; 96–98% clean scans.

Caveats

Diacritic-handling validated; bilingual EN/ES docs use dominant-script routing.

French

fr
Script
Latin
Engine
Tesseract
Accuracy
97–99% on born-digital; 95–97% clean scans.

Caveats

Diacritic-heavy text fully validated.

German

de
Script
Latin
Engine
Tesseract
Accuracy
97–99% on born-digital; 95–97% clean scans.

Caveats

Umlauts + ß handled; compound nouns parsed correctly.

Italian

it
Script
Latin
Engine
Tesseract
Accuracy
97–99% on born-digital.

Caveats

Accent handling validated.

Dutch

nl
Script
Latin
Engine
Tesseract
Accuracy
97–99% on born-digital.

Caveats

Compound-word handling validated.

Portuguese

pt
Script
Latin
Engine
Tesseract
Accuracy
95–97% on clean printed text.

Caveats

BR and PT variants both supported; currency/date formats locale-aware.

Chinese (Simplified)

zh-Hans
Script
Han
Engine
PaddleOCR
Accuracy
94–96% on clean printed text.

Caveats

Printed-only at Tier 2A; handwritten CJK routes to Gemini (HW-NE mode).

Chinese (Traditional)

zh-Hant
Script
Han
Engine
PaddleOCR
Accuracy
92–95% on clean printed text.

Caveats

Some less-frequent traditional characters benefit from HW-NE mode.

Japanese

ja
Script
Han + Hiragana + Katakana
Engine
PaddleOCR
Accuracy
93–96% on clean printed text.

Caveats

Handwritten Japanese (HW-NE mode) uses Gemini Flash — NLS 0.899 benchmark.

Korean

ko
Script
Hangul
Engine
PaddleOCR
Accuracy
92–95% on clean printed text.

Caveats

Mixed Hangul/Hanja documents (formal text) handled.

Hindi

hi
Script
Devanagari
Engine
Surya
Accuracy
90–94% on clean printed text.

Caveats

Handwritten Hindi: HW-NE mode (Qwen3-VL).

Marathi

mr
Script
Devanagari
Engine
Surya
Accuracy
88–92% on clean printed text.

Caveats

Shared Devanagari model with Hindi.

Tamil

ta
Script
Tamil
Engine
Surya
Accuracy
88–93% on clean printed text.

Caveats

Complex consonant clusters fully validated.

Telugu

te
Script
Telugu
Engine
Surya
Accuracy
88–93% on clean printed text.

Caveats

Compound vowel signs handled correctly.

Bengali

bn
Script
Bengali
Engine
Surya
Accuracy
88–92% on clean printed text.

Caveats

Conjuncts and ligatures handled.

Punjabi

pa
Script
Gurmukhi
Engine
Surya
Accuracy
88–92% on clean printed text.

Caveats

Gurmukhi-specific model; eMunshi corpus has validated 1M+ pages.

Arabic

ar
Script
Arabic (RTL)
Engine
Surya
Accuracy
89–93% on clean printed RTL text.

Caveats

Handwritten Arabic routes to GPT-4o (HW-NE mode); cursive ligatures are the hardest non-CJK problem.

Tier · Beta

Beta support

Works on clean documents. Real-world accuracy depends on your document quality. Validate on a sample before committing to volume.

Vietnamese

vi
Script
Latin (diacritic-heavy)
Engine
Surya
Accuracy
93–96% on clean printed text.

Caveats

Stacked tone marks; validate on your scan quality.

Tagalog

tl
Script
Latin
Engine
Tesseract
Accuracy
94–96% on clean printed text.

Caveats

Spanish loanwords and abbreviations handled correctly.

Nepali

ne
Script
Devanagari
Engine
Surya
Accuracy
85–90% on clean printed text.

Caveats

Validate on real corpus before production deployment.

Kannada

kn
Script
Kannada
Engine
Surya
Accuracy
85–90% on clean printed text.

Caveats

Less-common conjuncts may need HW-NE escalation.

Malayalam

ml
Script
Malayalam
Engine
Surya
Accuracy
85–90% on clean printed text.

Caveats

Atomic-character changes since 1971 reform — older documents may need HW-NE.

Odia

or
Script
Bengali (Odia variant)
Engine
Surya
Accuracy
83–88% on clean printed text.

Caveats

Shared model with Bengali; validate per use case.

Gujarati

gu
Script
Gujarati
Engine
Surya
Accuracy
85–90% on clean printed text.

Caveats

Validate on stamped/scanned government forms.

Persian

fa
Script
Arabic (RTL)
Engine
Surya
Accuracy
86–91% on clean printed RTL text.

Caveats

Shared model with Arabic + Persian-specific chars validated.

Urdu

ur
Script
Arabic (Nastaliq)
Engine
Surya
Accuracy
83–88% on clean printed Nastaliq text.

Caveats

Nastaliq is harder than naskh; cursive variants benefit from HW-NE.

Hebrew

he
Script
Hebrew (RTL)
Engine
Surya
Accuracy
85–90% on clean printed RTL text.

Caveats

Handwritten Hebrew → GPT-4o (no open-source handwriting model exists).

Thai

th
Script
Thai
Engine
Surya
Accuracy
85–90% on clean printed text.

Caveats

No word boundaries in Thai; OCR confidence calibration varies.

Russian

ru
Script
Cyrillic
Engine
Tesseract
Accuracy
95–97% on clean printed Cyrillic.

Caveats

Handwriting → Qwen3-VL.

Ukrainian

uk
Script
Cyrillic
Engine
Tesseract
Accuracy
94–96% on clean printed Cyrillic.

Caveats

Ukrainian-specific characters fully supported.

Bulgarian

bg
Script
Cyrillic
Engine
Tesseract
Accuracy
93–96% on clean printed Cyrillic.

Caveats

Validate on real corpus for production.

Greek

el
Script
Greek
Engine
Tesseract
Accuracy
93–96% on clean printed Greek.

Caveats

Modern Greek validated; polytonic ancient Greek experimental.

Tier · Experimental

Experimental support

Active development. Accuracy varies. The hardest scripts most tools fail on entirely — we ship them honestly labeled instead of overpromising.

Khmer

km
Script
Khmer
Engine
Surya
Accuracy
75–85% on clean printed text.

Caveats

Handwritten Khmer always flags for human review — no production-grade model exists.

Serbian

sr
Script
Cyrillic
Engine
Tesseract
Accuracy
90–93% on clean printed Cyrillic.

Caveats

Latin Serbian also supported via Latin pipeline.

How we set the tiers

Honest beats optimistic.

Stablemeans we have run thousands of synthetic and anonymized real-world documents through it, validated extraction accuracy against ground truth, and shipped it as production-ready. We'd use it ourselves in a regulated workflow.

Betameans the OCR engine and routing work, but the maturity isn't backed by exhaustive validation. We've tested it on clean documents. Your scan quality, document layout, and font choice might surface failures we haven't seen. Validate on a sample first.

Experimentalmeans it works enough to be useful, but accuracy varies meaningfully across documents. Indic scripts are the hardest cases in OCR — most tools refuse to ship them at all, or ship them with misleading "supported" labels. We ship them with honest expectations instead.

As accuracy improves, languages get promoted. Promotions are documented in the CHANGELOG with the validation work that supported them.

Language not on this list?

Inspire AI Lab has run extraction at scale on a 230M-document multilingual legal corpus. If you have a custom language requirement — fine-tuning, new script support, dialect handling — we can scope a custom build against your real corpus.