πŸ”¬ OlmOCR-Bench & Real-World Document Benchmark Suite (2026)

OCR API & VLM Accuracy Benchmarks

Deterministic pass/fail unit testing (8,413 tests across 1,403 complex PDF pages) by AllenAI (Ai2). Evaluates spatial neighbor relationships, LaTeX equations, multi-column reading orders, and single-page inference latency.

πŸ† OlmOCR-Bench Architectural Leaderboard (Overall)

Deterministic unit testing across 7 categories (ArXiv Math, Table Structure, Header/Footer suppression, Multi-Column, Historical Scans).

Model & Engine Architecture OlmOCR-Bench Score Key Differentiators
Chandra OCR SOTA
29B VLM 85.8 Multilingual (90+), checkbox reconstruction, HTML/JSON output
Mistral OCR 4
Proprietary API 85.20 Spatial bounding boxes, block classification, per-word confidence scores
LlamaParse (Agentic Plus)
Multi-tier VLM Routing 83.5 Cost Optimizer auto-routing, native chart & financial parsing
olmOCR-2
7B VLM (Qwen2.5-VL) 82.4 RLVR training, verifiable reward loops, perfect mathematical LaTeX
Reducto
Proprietary VLM 81.0 SOC2/HIPAA zero retention, clean Markdown & HTML parsing
PaddleOCR-VL
0.9B VLM (ERNIE 4.5) 80.0 NaViT visual encoder, dynamic aspect ratio, 109 languages
Azure Document Intelligence
Layout AI VLM 78.2 14 prebuilt schemas, coordinate geometry mapping
Google Cloud Document AI
Document AI VLM 77.0 BigQuery native warehouse routing, 200+ languages
AWS Textract
ML Vision Pipeline 76.5 AnalyzeExpense endpoint, spatial polygon frame blocks
DeepSeek-OCR
3B VLM (DeepSeek-MoE) 75.7 Contextual Optical Compression, 10x-20x fewer vision tokens
ABBYY FineReader Engine
Deterministic OCR + ML 74.0 Legacy gold standard for historical & degraded physical scans
Nanonets-OCR2-3B
4B VLM 69.5 Mermaid flowcharts, watermark & signature detection
Tesseract OCR
LSTM Pipeline 52.0 Clean text only, no layout or table understanding

πŸ“Š OlmOCR-Bench: Deterministic Unit-Test Score (Overall)

8,413 deterministic pass/fail unit tests evaluated across 1,403 complex PDF pages testing spatial neighbor relationships, LaTeX equations, and multi-column reading order.

Metric: OlmOCR-Bench Score (Higher is better - Max 100)
Rank Provider & Engine Benchmark Metric Score / Accuracy Action
πŸ₯‡ Mistral OCR 4 85.20
85.2%
Review
πŸ₯ˆ LlamaParse (Agentic) 83.50
83.5%
Review
πŸ₯‰ olmOCR-2 (7B VLM) 82.40
82.4%
Review
#4 Reducto AI 81.00
81%
Review
#5 PaddleOCR-VL (0.9B) 80.00
80%
Review
#6 Azure AI Doc Intelligence 78.20
78.2%
Review
#7 Google Cloud Doc AI 77.00
77%
Review
#8 AWS Textract 76.50
76.5%
Review
#9 DeepSeek-OCR (3B) 75.70
75.7%
Review
#10 ABBYY FineReader Engine 74.00
74%
Review
#11 Nanonets OCR 2 (3B) 69.50
69.5%
Review
#12 Tesseract OCR 52.00
52%
Review

πŸ“Š Complex Tables & Financial Statements (ICDAR / FinTabNet)

3,000 dense financial tables with merged cells, multi-line headers, and borderless columns.

Metric: TEDS Score (Tree-Edit-Distance Similarity % - Higher is better)
Rank Provider & Engine Benchmark Metric Score / Accuracy Action
πŸ₯‡ Mistral OCR 4 95.8%
95.8%
Review
πŸ₯ˆ olmOCR-2 (7B) 95.5%
95.5%
Review
πŸ₯‰ LlamaParse 95.2%
95.2%
Review
#4 Azure Doc Intelligence 94.5%
94.5%
Review
#5 PaddleOCR-VL (0.9B) 94.0%
94%
Review
#6 AWS Textract 93.8%
93.8%
Review
#7 Reducto AI 93.5%
93.5%
Review
#8 DeepSeek-OCR (3B) 93.0%
93%
Review
#9 Mindee 92.0%
92%
Review
#10 ABBYY FineReader Engine 91.2%
91.2%
Review
#11 Nanonets OCR 2 91.0%
91%
Review
#12 Google Cloud Doc AI 88.2%
88.2%
Review
#13 Tesseract OCR 61.2%
61.2%
Review

πŸ“Š Invoices & Receipts Key Information Extraction (SROIE / CORD)

3,500 skewed, wrinkled mobile photos and PDF invoices with multi-item tax breakdowns.

Metric: Key Information Extraction (F1 Score % - Higher is better)
Rank Provider & Engine Benchmark Metric Score / Accuracy Action
πŸ₯‡ Mistral OCR 4 95.8%
95.8%
Review
πŸ₯ˆ Mindee Invoice API 95.8%
95.8%
Review
πŸ₯‰ LlamaParse 95.0%
95%
Review
#4 olmOCR-2 (7B) 94.8%
94.8%
Review
#5 Azure Doc Intelligence 94.2%
94.2%
Review
#6 AWS Textract (AnalyzeExpense) 93.5%
93.5%
Review
#7 PaddleOCR-VL (0.9B) 93.2%
93.2%
Review
#8 Reducto AI 93.0%
93%
Review
#9 DeepSeek-OCR (3B) 92.0%
92%
Review
#10 ABBYY FineReader Engine 91.0%
91%
Review
#11 Google Cloud Doc AI 91.0%
91%
Review
#12 Tesseract OCR 68.5%
68.5%
Review

πŸ“Š Inference Latency & Throughput (Single A100 GPU / Cloud Endpoint)

Response time for single-page 300 DPI documents under 50 concurrent requests.

Metric: p50 Response Time (Milliseconds - Lower is faster)

Need reproduction scripts and dataset details?

Read our comprehensive test harness methodology and test data disclosure.

Read Methodology Document β†’

❓ OlmOCR-Bench & Benchmark Methodology FAQs

What is OlmOCR-Bench and why is it the 2026 standard? β–Ό

Published by AllenAI (Ai2), OlmOCR-Bench is the rigorous standard for evaluating Vision-Language OCR models. It abandons string-matching CER/WER in favor of 8,413 deterministic, pass/fail unit tests across 1,403 complex PDF pages testing ArXiv Math LaTeX rendering, spatial neighbor table cell relationships, multi-column reading orders, and historical scans.

Why does traditional Edit Distance fail on LaTeX equations? β–Ό

Traditional Edit Distance heavily penalizes a model that outputs \dfrac{a}{b} instead of a ground truth \frac{a}{b}, even though both render the exact same mathematical formula in LaTeX. OlmOCR-Bench evaluates the rendered DOM element bounding boxes, correctly scoring mathematically equivalent outputs as successful passes.

What is the TEDS metric for Table Extraction? β–Ό

Tree-Edit-Distance-based Similarity (TEDS) evaluates both cell text accuracy and the structural tree hierarchy (row spans, column spans, headers) of HTML/Markdown tables. A TEDS score above 92% is considered production-ready for financial automation.

How do open-source models like olmOCR-2 compare to closed APIs? β–Ό

AllenAI’s olmOCR-2 (82.4 score on OlmOCR-Bench) and PaddleOCR-VL-0.9B (80.0 score) now outperform several proprietary commercial APIs on complex mathematical typography and multi-column layouts, while costing ~$176 in GPU compute per 1,000,000 pages.