PAVE: Learning Physics-Aware Visual Representations from Videos
Abstract
Physical knowledge is essential for physical reasoning, video generation, and robotic manipulation. However, general-purpose visual encoders primarily learn semantic and spatial features, with limited emphasis on physical properties and interactions. We introduce PAVE, a vision encoder that learns transferable physical representations from videos. For data, we curate PhysCentric, a corpus of 515,377 unlabeled video clips covering diverse physical interactions. For architecture, we combine sparsely routed experts with a shared expert to capture specialized and common physical cues. For training, we first learn physical dynamics through causal future-feature prediction, then apply Physics Knowledge Alignment using physical question-answer supervision and Baseline-Adjusted Cross-Entropy (BACE) to support physical reasoning. PAVE achieves 58.93% on MVBench and 70.04% on PhysBench, surpassing the strongest compared encoders by 3.08 and 2.01 percentage points, respectively. Downstream gains over the respective baselines include 9.7 percentage points in VideoPhy semantic adherence, 6.4 points in VSI-Bench average score, and 3.2 percentage points in average success rate across five RoboTwin 2.0 tasks, which shows the strong transferability. The data, code, and model will be publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.