OCR API & VLM Accuracy Benchmarks
Deterministic pass/fail unit testing (8,413 tests across 1,403 complex PDF pages) by AllenAI (Ai2). Evaluates spatial neighbor relationships, LaTeX equations, multi-column reading orders, and single-page inference latency.
π OlmOCR-Bench Architectural Leaderboard (Overall)
Deterministic unit testing across 7 categories (ArXiv Math, Table Structure, Header/Footer suppression, Multi-Column, Historical Scans).
| Model & Engine | Architecture | OlmOCR-Bench Score | Key Differentiators |
|---|---|---|---|
| Chandra OCR SOTA | 29B VLM | 85.8 | Multilingual (90+), checkbox reconstruction, HTML/JSON output |
| Mistral OCR 4 | Proprietary API | 85.20 | Spatial bounding boxes, block classification, per-word confidence scores |
| LlamaParse (Agentic Plus) | Multi-tier VLM Routing | 83.5 | Cost Optimizer auto-routing, native chart & financial parsing |
| olmOCR-2 | 7B VLM (Qwen2.5-VL) | 82.4 | RLVR training, verifiable reward loops, perfect mathematical LaTeX |
| Reducto | Proprietary VLM | 81.0 | SOC2/HIPAA zero retention, clean Markdown & HTML parsing |
| PaddleOCR-VL | 0.9B VLM (ERNIE 4.5) | 80.0 | NaViT visual encoder, dynamic aspect ratio, 109 languages |
| Azure Document Intelligence | Layout AI VLM | 78.2 | 14 prebuilt schemas, coordinate geometry mapping |
| Google Cloud Document AI | Document AI VLM | 77.0 | BigQuery native warehouse routing, 200+ languages |
| AWS Textract | ML Vision Pipeline | 76.5 | AnalyzeExpense endpoint, spatial polygon frame blocks |
| DeepSeek-OCR | 3B VLM (DeepSeek-MoE) | 75.7 | Contextual Optical Compression, 10x-20x fewer vision tokens |
| ABBYY FineReader Engine | Deterministic OCR + ML | 74.0 | Legacy gold standard for historical & degraded physical scans |
| Nanonets-OCR2-3B | 4B VLM | 69.5 | Mermaid flowcharts, watermark & signature detection |
| Tesseract OCR | LSTM Pipeline | 52.0 | Clean text only, no layout or table understanding |
π OlmOCR-Bench: Deterministic Unit-Test Score (Overall)
8,413 deterministic pass/fail unit tests evaluated across 1,403 complex PDF pages testing spatial neighbor relationships, LaTeX equations, and multi-column reading order.
| Rank | Provider & Engine | Benchmark Metric | Score / Accuracy | Action |
|---|---|---|---|---|
| π₯ | Mistral OCR 4 | 85.20 | 85.2% | Review |
| π₯ | LlamaParse (Agentic) | 83.50 | 83.5% | Review |
| π₯ | olmOCR-2 (7B VLM) | 82.40 | 82.4% | Review |
| #4 | Reducto AI | 81.00 | 81% | Review |
| #5 | PaddleOCR-VL (0.9B) | 80.00 | 80% | Review |
| #6 | Azure AI Doc Intelligence | 78.20 | 78.2% | Review |
| #7 | Google Cloud Doc AI | 77.00 | 77% | Review |
| #8 | AWS Textract | 76.50 | 76.5% | Review |
| #9 | DeepSeek-OCR (3B) | 75.70 | 75.7% | Review |
| #10 | ABBYY FineReader Engine | 74.00 | 74% | Review |
| #11 | Nanonets OCR 2 (3B) | 69.50 | 69.5% | Review |
| #12 | Tesseract OCR | 52.00 | 52% | Review |
π Complex Tables & Financial Statements (ICDAR / FinTabNet)
3,000 dense financial tables with merged cells, multi-line headers, and borderless columns.
| Rank | Provider & Engine | Benchmark Metric | Score / Accuracy | Action |
|---|---|---|---|---|
| π₯ | Mistral OCR 4 | 95.8% | 95.8% | Review |
| π₯ | olmOCR-2 (7B) | 95.5% | 95.5% | Review |
| π₯ | LlamaParse | 95.2% | 95.2% | Review |
| #4 | Azure Doc Intelligence | 94.5% | 94.5% | Review |
| #5 | PaddleOCR-VL (0.9B) | 94.0% | 94% | Review |
| #6 | AWS Textract | 93.8% | 93.8% | Review |
| #7 | Reducto AI | 93.5% | 93.5% | Review |
| #8 | DeepSeek-OCR (3B) | 93.0% | 93% | Review |
| #9 | Mindee | 92.0% | 92% | Review |
| #10 | ABBYY FineReader Engine | 91.2% | 91.2% | Review |
| #11 | Nanonets OCR 2 | 91.0% | 91% | Review |
| #12 | Google Cloud Doc AI | 88.2% | 88.2% | Review |
| #13 | Tesseract OCR | 61.2% | 61.2% | Review |
π Invoices & Receipts Key Information Extraction (SROIE / CORD)
3,500 skewed, wrinkled mobile photos and PDF invoices with multi-item tax breakdowns.
| Rank | Provider & Engine | Benchmark Metric | Score / Accuracy | Action |
|---|---|---|---|---|
| π₯ | Mistral OCR 4 | 95.8% | 95.8% | Review |
| π₯ | Mindee Invoice API | 95.8% | 95.8% | Review |
| π₯ | LlamaParse | 95.0% | 95% | Review |
| #4 | olmOCR-2 (7B) | 94.8% | 94.8% | Review |
| #5 | Azure Doc Intelligence | 94.2% | 94.2% | Review |
| #6 | AWS Textract (AnalyzeExpense) | 93.5% | 93.5% | Review |
| #7 | PaddleOCR-VL (0.9B) | 93.2% | 93.2% | Review |
| #8 | Reducto AI | 93.0% | 93% | Review |
| #9 | DeepSeek-OCR (3B) | 92.0% | 92% | Review |
| #10 | ABBYY FineReader Engine | 91.0% | 91% | Review |
| #11 | Google Cloud Doc AI | 91.0% | 91% | Review |
| #12 | Tesseract OCR | 68.5% | 68.5% | Review |
π Inference Latency & Throughput (Single A100 GPU / Cloud Endpoint)
Response time for single-page 300 DPI documents under 50 concurrent requests.
| Rank | Provider & Engine | Benchmark Metric | Score / Accuracy | Action |
|---|---|---|---|---|
| π₯ | PaddleOCR-VL (Local Edge/GPU) | 110 ms | 99% | Review |
| π₯ | DeepSeek-OCR (vLLM A100) | 120 ms | 98% | Review |
| π₯ | Mindee Cloud API | 350 ms | 95% | Review |
| #4 | Nanonets OCR 2 (Local GPU) | 380 ms | 92% | Review |
| #5 | olmOCR-2 (vLLM A100) | 420 ms | 90% | Review |
| #6 | Tesseract OCR (Local CPU) | 420 ms | 90% | Review |
| #7 | Reducto AI Cloud | 480 ms | 88% | Review |
| #8 | Mistral OCR 4 Cloud | 550 ms | 85% | Review |
| #9 | Google Cloud Doc AI | 680 ms | 80% | Review |
| #10 | Azure AI Doc Intelligence | 720 ms | 78% | Review |
| #11 | AWS Textract | 850 ms | 72% | Review |
| #12 | LlamaParse Cloud | 950 ms | 65% | Review |
| #13 | ABBYY FineReader Engine | 1,400 ms | 50% | Review |
Need reproduction scripts and dataset details?
Read our comprehensive test harness methodology and test data disclosure.
β OlmOCR-Bench & Benchmark Methodology FAQs
What is OlmOCR-Bench and why is it the 2026 standard? βΌ
Published by AllenAI (Ai2), OlmOCR-Bench is the rigorous standard for evaluating Vision-Language OCR models. It abandons string-matching CER/WER in favor of 8,413 deterministic, pass/fail unit tests across 1,403 complex PDF pages testing ArXiv Math LaTeX rendering, spatial neighbor table cell relationships, multi-column reading orders, and historical scans.
Why does traditional Edit Distance fail on LaTeX equations? βΌ
Traditional Edit Distance heavily penalizes a model that outputs \dfrac{a}{b} instead of a ground truth \frac{a}{b}, even though both render the exact same mathematical formula in LaTeX. OlmOCR-Bench evaluates the rendered DOM element bounding boxes, correctly scoring mathematically equivalent outputs as successful passes.
What is the TEDS metric for Table Extraction? βΌ
Tree-Edit-Distance-based Similarity (TEDS) evaluates both cell text accuracy and the structural tree hierarchy (row spans, column spans, headers) of HTML/Markdown tables. A TEDS score above 92% is considered production-ready for financial automation.
How do open-source models like olmOCR-2 compare to closed APIs? βΌ
AllenAIβs olmOCR-2 (82.4 score on OlmOCR-Bench) and PaddleOCR-VL-0.9B (80.0 score) now outperform several proprietary commercial APIs on complex mathematical typography and multi-column layouts, while costing ~$176 in GPU compute per 1,000,000 pages.