Does Calibrated Surprise Detect Physical Violations? Auditing Conformal Monitors for Latent World Models
Abstract
Prediction error in a latent world model, or surprise, can flag an unexpected transition. What makes that alarm evidence of a physical inconsistency? We introduce VoE-Bench, a violation-of-expectation benchmark combining an audit of 29 checkpoints in three environments with focused tests of reference sensitivity, physical selectivity, and decision utility. In TwoRoom, changing the clean (unedited, default-rendering) reference moves alarms on the same 1,000 benignly recoloured episodes from to , although both nominal audits meet tolerance. Fresh references (9,600 episodes) confirm this sensitivity and reveal discrepancies from a point forecast frozen before scoring. In PushT, an action-aware multilayer perceptron (MLP) detects more action mismatches than its action-free counterpart under separate calibration. At approximately empirical nominal false-positive rate, the original head has no resolved advantage over the action-aware MLP sharing its encoder and inputs in three-frame onset-window detection ( percentage points; interval ). Under constant recolouring, mean teleport ordering above an appearance-only twin is when the teleported object is recoloured and when the other object is (other-minus-same difference ; interval ). TwoRoom readouts recover position despite unselective residuals, and later PushT residuals reveal action evidence masked by an appearance transient. Separate planning tests resolve no success gain and expose the cost of mismatched screening references. The audit distinguishes reference sensitivity, appearance-dependent ordering, and temporal masking, identifying what must be checked before interpreting or acting on an alarm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.