acceptodds
Under review as a conference paper at ICLR 2027

Does Surprise Measure Intuitive Physics? The Score Depends on How the Model Is Read

Abstract

Latent-predictive video models are said to acquire intuitive physics when they are more surprised by an impossible event than by a possible one (Garrido et al., 2025). That surprise is a single scalar, the difference in prediction error between two otherwise identical clips. The benchmark that defines that score chose the measure deliberately, naming one alternative protocol (Weihs et al., 2022) and reporting that the alternative had not yielded evidence of understanding (Bordes et al., 2025). Nothing has established what the difference contains. We decompose the squared form exactly into magnitude, direction and a selection term, and show it is not a fixed property of the model. We measure it on V-JEPA 2 (Assran et al., 2025) and V-JEPA 2.1 (Mur-Labadia et al., 2026). Both are publicly released latent-predictive encoders, and only V-JEPA 2.1 exposes intermediate depths to read at. Reading V-JEPA 2.1 at the four encoder depths its predictor was trained to read, with nothing attached on top, the reported accuracy spans 10.18 percentage points with the model, the clips and the readout all unchanged. Across the two encoders every term keeps its sign and none keeps its size, so the two standard normalisations disagree about which model is more direction-driven. A motion-only observer touching neither model shows that most or more of what the score detects is available from movement alone. We then move the score using a regulariser carrying no physics. An untrained random projection preserves accuracy at 88.39% against the frozen encoder's 89.55%. Training a small network in its place, under the coefficients published with VICReg (Bardes et al., 2022), drives it to 46.57%, a point estimate below chance whose 95% confidence interval contains chance. Raising one of those three coefficients recovers 94.62%. A violation-of-expectation score is not readable as physics evidence until six quantities are reported beside it. We specify those six and show what each one catches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.