Physics Before the First Step: Rectifying the Text Embedding of Frozen Video Generators
Abstract
Video generators trained at scale acquire a prior over natural motion, yet they still fail on actions whose outcome is dictated by physics. We observe that as flow matching and DiT backbones replace earlier designs, the structure of the generated video is decided earlier and earlier in the denoising process. This undermines methods that correct the latent during sampling, which need a window in which an error is visible and still open to repair. We turn instead to the text embedding, the one input read clean at every layer and every step, and find two properties that make it a sufficient and learnable target: an embedding fitted to a single real clip reproduces that clip under every initial noise, and the displacement to it is spread across the whole embedding, so a prediction that lands near it suffices. Building on these, we propose PhysRect: for physical events cut from real video, it fits the embedding under which a frozen generator reproduces the clip, keeping the fits of different mechanisms apart with a separability term. The fits are then distilled into a VLM-gated mixture of mechanism experts that corrects the embedding once before sampling. On VideoPhy-2, PhysRect reduces the physics error of Wan2.2 by 51%, matches a 14B model fine-tuned on physical video at one fourteenth of its generation time with no inference overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.