acceptodds
Under review as a conference paper at ICLR 2027

PAVE: Learning Physics-Aware Visual Representations from Videos

Abstract

Physical knowledge is essential for physical reasoning, video generation, and robotic manipulation. However, general-purpose visual encoders primarily learn semantic and spatial features, with limited emphasis on physical properties and interactions. We introduce PAVE, a vision encoder that learns transferable physical representations from videos. For data, we curate PhysCentric, a corpus of 515,377 unlabeled video clips covering diverse physical interactions. For architecture, we combine sparsely routed experts with a shared expert to capture specialized and common physical cues. For training, we first learn physical dynamics through causal future-feature prediction, then apply Physics Knowledge Alignment using physical question-answer supervision and Baseline-Adjusted Cross-Entropy (BACE) to support physical reasoning. PAVE achieves 58.93% on MVBench and 70.04% on PhysBench, surpassing the strongest compared encoders by 3.08 and 2.01 percentage points, respectively. Downstream gains over the respective baselines include 9.7 percentage points in VideoPhy semantic adherence, 6.4 points in VSI-Bench average score, and 3.2 percentage points in average success rate across five RoboTwin 2.0 tasks, which shows the strong transferability. The data, code, and model will be publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.