Where our numbers come from
Every benchmark figure on VoidSource is measured, curated, or tracked, and labeled as such.
Measured means we ran the evaluation ourselves and archived the raw output. Curated means we digitized a table verbatim from a named source and dated it. Tracked means a public leaderboard flows into our model pages through the daily pipeline. Nothing on this site mixes the three without saying so.
In-house runs with the official evaluator, raw predictions archived
olmOCR-Bench reproduction
last verified 2026-05-22olmOCR-bench scores reported by VoidSource are reproduced in-house with the official olmocr.bench evaluator on the allenai/olmOCR-bench 1.0 snapshot (1,403 PDFs); per-model raw predictions are archived on Cloudflare R2. Headline figures last verified 2026-05-22 against the run index dated 2026-03-21.
- Models run
- 15
- Dataset
- allenai/olmOCR-bench @ 1.0
- Run index
- voidsourceData/packages/vsd-bench/results/olmocr-bench-results.json
- Raw archive
- experiments/olmocr/{model}/{version}.tar.gz
Upstream results digitized verbatim, dated, and traced to their source
olmOCR-Bench
curated 2026-01-22Unit-test style evaluation for PDF linearization quality. Digitized from Allen AI olmOCR GitHub, HuggingFace Dataset, LightOnOCR arXiv Paper, not measured by us.
Top 5
OmniDocBench v1.5
curated 2026-01-28Comprehensive document reading evaluation across text, formulas, tables, and reading order. Digitized from DeepSeek-OCR-2 Paper, DeepSeek-OCR-2 HuggingFace, OmniDocBench GitHub, not measured by us.
Top 5
Public leaderboards the daily pipeline joins into the model pages
17 leaderboards, joined across 272 model variants. These are not our measurements: each score belongs to its leaderboard and is re-scraped daily, so the model pages never show a number older than the source. Scores appear in context on the LLM hub and the per-model pages.
last pipeline join: 2026-07-25