acceptodds
Under review as a conference paper at ICLR 2027

Knowing When Not to Answer: Self-Assessed Measurability and the Limits of Physics Transfer

Abstract

A world model that has learned physics should carry something to physics it has not seen, and the field tests this by training on one system and evaluating on another. We show that this test usually does not measure what it is taken to. The reason is structural: a latent world model compresses a field, evolves it and decompresses, so its error contains both what the dynamics got wrong and what the round trip lost, and which term dominates depends on something rarely reported, namely how far the physics moves between the two frames being compared. When the field barely moves and the codec is lossy, the error measures the codec, and a model that has learned nothing transferable still scores well. The fix is a check the model runs on itself before acting, with no labels and no access to the future. Let ρ be its reconstruction error divided by the observed frame-to-frame change: below 1 the comparison is attributable to dynamics, at or above 1 it is not. The threshold is derived rather than fitted, and over 52 conditions the rule makes no errors; two stress tests and two weaker checks leave it intact. Applied to a standard PDE benchmark it disqualifies the protocol rather than the models: of 42 cross-family pairs across eight datasets, none is measurable at the single-step horizon those benchmarks use. This is not undertraining, nor simply a weak codec. Training the source to convergence improves in-domain attributability by over an order of magnitude and cross-domain attributability by under 2×, while refitting the codec on the target buys attributability only by destroying the skill being measured. What survives is narrower and real. Inside one physics family transfer occurs and is measurable, in three of six directions among three acoustic geometries, one beating codec-matched persistence by 0.568 [0.530, 0.606] at every horizon and replicating across seeds; across families, nowhere. The consequence is a world model that knows when not to answer, whose failure mode is decidable rather than probabilistic: we placed conditions on the boundary itself and the classes still did not overlap. It also predicts the size of its own forecast error well enough to be calibrated, at 90.6% coverage for a nominal 90% once the interval accounts for shift between families.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.