Arabic OCR

Benchmark published · vendor-conducted, methodology open

We didn’t claim Arabic OCR leadership. We measured it.

Evaluated against 11 leading OCR and AI systems across 3,760 documents. YaiGlobal ranked #1 of 12 on both accuracy metrics that matter.

0.125% CER 0.250% WER #1 of 12 systems evaluated
scanned_page.tif

تقرير الأرشيف الرقمي — القسم الثاني

تتيح تقنيات التعرف الضوئي الحديثة استخلاص النصوص العربية بدقة عالية مع الحفاظ الكامل على بنية الصفحة الأصلية وترتيب القراءة.

output.json
"language": "ar", "direction": "rtl", "cer": 0.00125, "wer": 0.00250, "structures": ["heading", "paragraph", "table"]

The proof

Arabic depth you can measure — not just a claim.

3,760 documents. 12 benchmark datasets. 11 competing systems, including Google, OpenAI and Microsoft. YaiGlobal ranked #1 overall on both accuracy metrics.

  • 0.250%Average Word Error Rate — best of 12 systems
  • 0.125%Average Character Error Rate — best of 12 systems
  • 3,760Documents evaluated across all datasets
  • 12Datasets — ranked #1 overall on both metrics
Average Word Error Rate by system Lower is better · 3,760 documents

Average Word Error Rate, lower is better. YaiGlobal also ranked #1 on Character Error Rate. Tesseract, Paddle, Qwen2.5-VL, Qwen2-VL and Surya scored higher error than every system shown here.

YaiGlobal

Avg. WER0.250%

Avg. CER0.125%

Datasets won on WER7 / 12

Gemini 2.0 Flash — Google

Avg. WER0.313%

Avg. CER0.128%

Datasets won on WER5 / 12

44%

Lower WER than Qari

63%

Lower CER than Qari

58%

Lower WER than AIN

69%

Lower CER than AIN

Why Arabic OCR is different

Arabic isn’t just another OCR problem.

Every item below comes from the benchmark’s own framing and from real production deployments — not a marketing checklist.

  • RTL reading order

    Right-to-left flow that must stay correct even inside mixed-language lines and tables.

  • Positional letterforms

    Letters change shape by position in a word — initial, medial, final, isolated.

  • Optional diacritics

    Vowel marks (تشكيل) that are often present, often absent, always meaning-bearing.

  • Mixed Arabic / Latin

    Citations, model numbers and bidirectional lines combining both scripts in one sentence.

  • Historical & degraded print

    Faded ink, foxing, centuries-old typesetting — the real condition of archival collections.

  • Handwriting & calligraphy

    From everyday handwriting to ornate calligraphic manuscript hands.

  • Complex tables & columns

    Merged and nested tables, multi-column pages with embedded figures.

  • Bidirectional structure

    Headings, captions and notation that shift direction mid-page and must resolve correctly.

We built YaiGlobal by solving these problems first — not by adding Arabic support later.

See it happen

From a scanned page, out of the dark, into structured knowledge.

A single illustrative page, watched through YaiGlobal’s pipeline — region detection, reading order, recognition and structure, in sequence.

Plays while in view · pick a stage to jump

Scanned document

Heading1

حفظ التراث المخطوط العربي

Text2

تشهد المكتبات الرقمية اليوم تحولاً جذرياً في طرق حفظ الوثائق العربية ومعالجتها، إذ تتيح تقنيات التعرف الضوئي الحديثة استخلاص النصوص بدقة عالية.

Table3
القسمالصفحاتالحالة
الفصل الأول١–٤٨مكتمل
الفصل الثاني٤٩–١١٢قيد المراجعة
الملاحق١١٣–١٣٠مكتمل

Structured output

Illustrative page composed for this demo, not a customer document. Labels reflect structure detected — not per-field accuracy scores.

Built for Arabic

Arabic isn’t an afterthought at YaiGlobal.

A native Arabic-speaking team, working from two offices, built the engine that leads this benchmark — Arabic first, everything else extending from it.

  • Arabic-native intelligence

    Optimized and tested on the hardest Arabic documents — connected letterforms, calligraphy, degraded historical print.

  • Visual grounding

    Every extracted element traces back to the exact region of the page it came from.

  • Structural intelligence

    Headings, tables and reading order preserved — not just a stream of recognized characters.

  • Two offices, one product

    Santa Clara, USA and Ariana, Tunisia — engineering and native-language depth in the same team.

Beyond Arabic

Ready for every language — architecturally, not by accident.

The same architecture that leads on Arabic extends to multilingual and mixed-script documents. Arabic is the proof point and remains the benchmark leader; this section stays deliberately brief because we don’t publish numbers we haven’t measured yet.

From OCR to knowledge

Document → structure → search → knowledge.

Proven in production, not just in a benchmark: two real digitization engagements, presented at MELA 2026, Harvard University.

New York University — Arabic Collections Online

~1,000 older, degraded Arabic books, delivered as searchable hOCR

Real degraded historical print in active library production — a different and harder condition than the curated benchmark above.

1.2%

CER

3.0%

WER

0.9

Sec / page

Dar Almandumah

Structure-first pipeline, five coordinated AI agents

Orchestrator → Language Detection → Chunk Router → Specialist Agents → Layout & Assembly. Tables, headings and multi-column layout reconstructed, not flattened.

93.4%

Table fidelity (TEDS)

97.1%

Heading agreement

420

Pages / hour

Production case-study figures are a different, harder measurement condition than the controlled 3,760-document benchmark above (0.125% CER / 0.250% WER) — shown separately, never merged.

Deploy your way

Cloud, or fully on your infrastructure.

Web / Cloud

Start in minutes

The fastest path to production: the web platform and API, elastic and always up to date.

  • Online OCR platform and REST API
  • Elastic, multi-zone architecture
  • No infrastructure to manage

On-premise

Your infrastructure, your control

For organizations where sovereignty, privacy or compliance make this the only option.

  • Deployed entirely inside your environment
  • Data never leaves your infrastructure
  • Same engine, same accuracy, your terms

Custom solutions

Built for organizations with serious archives.

Libraries, universities, government, publishers, research institutions and enterprises with large Arabic document collections.

  • Digital library creation

    From scanned collection to structured, searchable library.

  • Search & discovery

    Citation linking, full-text search and retrieval-augmented Q&A over your collection.

  • Historical & institutional archives

    Manuscripts, degraded print and institutional repositories.

  • Custom integration

    API and workflow-specific tuning for existing systems.

Have Arabic documents?
Let’s benchmark them.

Send a sample. We’ll run it through the same pipeline that scored #1 of 12 in this report.