acceptodds
Under review as a conference paper at ICLR 2027

Steering Physics in Video World Models

Abstract

Video world models learn representations that support prediction and planning, and a growing body of work finds that these representations capture physics and dynamics. For example, velocity and acceleration can be read from the representations of V-JEPA 2. However, knowing that a quantity is encoded does not tell us where it is represented, and without that we cannot steer it. In this work, our goal is to edit the representation so that the decoded video carries out a requested physical instruction, such as making a ball move faster. Using simulation datasets, we investigate velocity and acceleration, both linear and angular. We first observe that translation and rotation trace very distinct shapes in the latent space, and that rotation splits along two axes. This structure motivates two steering writers, an additive writer that steers translation, and an orientation field that steers rotation. Measured from the motion in the decoded video, the additive writer controls both the direction and the speed of translation, with direction errors below and speed correlations of in 2D and in 3D. For rotation, the orientation field sets how fast the object spins, and the requested and decoded spin rates correlate at , holding at when the object's appearance changes. The same field, fitted only on objects spinning at a constant rate, also makes objects speed up or slow down their spin, with no refitting. Finally, we use an edited representation to steer a simulated robot arm, which strikes a ball within of the requested speed on all 200 test commands. Together, these results show that a video model's latent space can serve as a two-way interface to physics, in which physical quantities can be measured, set, and used as goals to plan actions in the real world.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.