YaiGlobal
Avg. WER0.250%
Avg. CER0.125%
Datasets won on WER7 / 12
Benchmark published · vendor-conducted, methodology open
Evaluated against 11 leading OCR and AI systems across 3,760 documents. YaiGlobal ranked #1 of 12 on both accuracy metrics that matter.
0.125% CER 0.250% WER #1 of 12 systems evaluatedتتيح تقنيات التعرف الضوئي الحديثة استخلاص النصوص العربية بدقة عالية مع الحفاظ الكامل على بنية الصفحة الأصلية وترتيب القراءة.
The proof
3,760 documents. 12 benchmark datasets. 11 competing systems, including Google, OpenAI and Microsoft. YaiGlobal ranked #1 overall on both accuracy metrics.
Average Word Error Rate, lower is better. YaiGlobal also ranked #1 on Character Error Rate. Tesseract, Paddle, Qwen2.5-VL, Qwen2-VL and Surya scored higher error than every system shown here.
YaiGlobal
Avg. WER0.250%
Avg. CER0.125%
Datasets won on WER7 / 12
Gemini 2.0 Flash — Google
Avg. WER0.313%
Avg. CER0.128%
Datasets won on WER5 / 12
44%
Lower WER than Qari
63%
Lower CER than Qari
58%
Lower WER than AIN
69%
Lower CER than AIN
Why Arabic OCR is different
Every item below comes from the benchmark’s own framing and from real production deployments — not a marketing checklist.
Right-to-left flow that must stay correct even inside mixed-language lines and tables.
Letters change shape by position in a word — initial, medial, final, isolated.
Vowel marks (تشكيل) that are often present, often absent, always meaning-bearing.
Citations, model numbers and bidirectional lines combining both scripts in one sentence.
Faded ink, foxing, centuries-old typesetting — the real condition of archival collections.
From everyday handwriting to ornate calligraphic manuscript hands.
Merged and nested tables, multi-column pages with embedded figures.
Headings, captions and notation that shift direction mid-page and must resolve correctly.
We built YaiGlobal by solving these problems first — not by adding Arabic support later.
See it happen
A single illustrative page, watched through YaiGlobal’s pipeline — region detection, reading order, recognition and structure, in sequence.
Scanned document
تشهد المكتبات الرقمية اليوم تحولاً جذرياً في طرق حفظ الوثائق العربية ومعالجتها، إذ تتيح تقنيات التعرف الضوئي الحديثة استخلاص النصوص بدقة عالية.
| القسم | الصفحات | الحالة |
|---|---|---|
| الفصل الأول | ١–٤٨ | مكتمل |
| الفصل الثاني | ٤٩–١١٢ | قيد المراجعة |
| الملاحق | ١١٣–١٣٠ | مكتمل |
Structured output
Illustrative page composed for this demo, not a customer document. Labels reflect structure detected — not per-field accuracy scores.
Built for Arabic
A native Arabic-speaking team, working from two offices, built the engine that leads this benchmark — Arabic first, everything else extending from it.
Optimized and tested on the hardest Arabic documents — connected letterforms, calligraphy, degraded historical print.
Every extracted element traces back to the exact region of the page it came from.
Headings, tables and reading order preserved — not just a stream of recognized characters.
Santa Clara, USA and Ariana, Tunisia — engineering and native-language depth in the same team.
Beyond Arabic
The same architecture that leads on Arabic extends to multilingual and mixed-script documents. Arabic is the proof point and remains the benchmark leader; this section stays deliberately brief because we don’t publish numbers we haven’t measured yet.
From OCR to knowledge
Proven in production, not just in a benchmark: two real digitization engagements, presented at MELA 2026, Harvard University.
New York University — Arabic Collections Online
Real degraded historical print in active library production — a different and harder condition than the curated benchmark above.
1.2%
CER
3.0%
WER
0.9
Sec / page
Dar Almandumah
Orchestrator → Language Detection → Chunk Router → Specialist Agents → Layout & Assembly. Tables, headings and multi-column layout reconstructed, not flattened.
93.4%
Table fidelity (TEDS)
97.1%
Heading agreement
420
Pages / hour
Production case-study figures are a different, harder measurement condition than the controlled 3,760-document benchmark above (0.125% CER / 0.250% WER) — shown separately, never merged.
Deploy your way
Web / Cloud
The fastest path to production: the web platform and API, elastic and always up to date.
On-premise
For organizations where sovereignty, privacy or compliance make this the only option.
Custom solutions
Libraries, universities, government, publishers, research institutions and enterprises with large Arabic document collections.
From scanned collection to structured, searchable library.
Citation linking, full-text search and retrieval-augmented Q&A over your collection.
Manuscripts, degraded print and institutional repositories.
API and workflow-specific tuning for existing systems.
Send a sample. We’ll run it through the same pipeline that scored #1 of 12 in this report.