Your Video Generator Secretly Contains Physics Experts
Abstract
Video generation models can produce high-quality, realistic videos, yet they often violate basic physical laws. Existing work largely treats these failures as evidence of missing capabilities that must be added through physical supervision, simulators, or additional training. In this paper, we find video generators capable of producing physically better videos already exist within pretrained video models. We can find those physics experts by applying random masks in the video models. Therefore, the physical accuracy of video generators can be improved without training. Across two model families, we also observe that at a fixed masking rate, these physics experts become denser as model scale increases. In practice, they can be selected without ground-truth videos, reused across scenes and obtained with comparable visual quality. Our results offer a new perspective on physical failures in video generation: physical capabilities may not be absent from pretrained models, but rather distributed across the diverse subnetworks they contain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.