Evidence

Where our numbers come from

Every benchmark figure on VoidSource is measured, curated, or tracked, and labeled as such.

Measured means we ran the evaluation ourselves and archived the raw output. Curated means we digitized a table verbatim from a named source and dated it. Tracked means a public leaderboard flows into our model pages through the daily pipeline. Nothing on this site mixes the three without saying so.

Measured by us

In-house runs with the official evaluator, raw predictions archived

olmOCR-Bench reproduction

last verified 2026-05-22

olmOCR-bench scores reported by VoidSource are reproduced in-house with the official olmocr.bench evaluator on the allenai/olmOCR-bench 1.0 snapshot (1,403 PDFs); per-model raw predictions are archived on Cloudflare R2. Headline figures last verified 2026-05-22 against the run index dated 2026-03-21.

Canary: olmOCR-2-7B-1025-FP8 reproduced 0.822 vs official leaderboard 0.824 (within stated ±1.1 tolerance) using the official evaluator.
Models run
15
Dataset
allenai/olmOCR-bench @ 1.0
Run index
voidsourceData/packages/vsd-bench/results/olmocr-bench-results.json
Raw archive
experiments/olmocr/{model}/{version}.tar.gz
Curated reference tables

Upstream results digitized verbatim, dated, and traced to their source

olmOCR-Bench

curated 2026-01-22

Unit-test style evaluation for PDF linearization quality. Digitized from Allen AI olmOCR GitHub, HuggingFace Dataset, LightOnOCR arXiv Paper, not measured by us.

Top 5

1.LightOnOCR-2-1B83.2%
2.Chandra OCR 0.1.083.1%
3.Infinity-Parser 7B82.5%
4.olmOCR v0.4.082.4%
5.LightOnOCR-2-1B-ocr-soup82.4%

OmniDocBench v1.5

curated 2026-01-28

Comprehensive document reading evaluation across text, formulas, tables, and reading order. Digitized from DeepSeek-OCR-2 Paper, DeepSeek-OCR-2 HuggingFace, OmniDocBench GitHub, not measured by us.

Top 5

1.PaddleOCR-VL92.86
2.DeepSeek-OCR 291.09
3.MinerU2.590.67
4.Qwen3-VL-235B89.15
5.MonkeyOCR-pro-3B88.85
Tracked leaderboards

Public leaderboards the daily pipeline joins into the model pages

17 leaderboards, joined across 272 model variants. These are not our measurements: each score belongs to its leaderboard and is re-scraped daily, so the model pages never show a number older than the source. Scores appear in context on the LLM hub and the per-model pages.

Aider PolyglotAIME 2025ARC-AGI-2GPQA DiamondGSOHumanity's Last ExamLiveBenchLiveCodeBenchLMArena EloMCP AtlasMMLU-ProSimpleBenchSWE-bench Pro (private set)SWE-bench Pro (public set)SWE-bench VerifiedTerminal-Benchτ-bench

last pipeline join: 2026-07-25