Condition-Dependent Probe Calibration in a World Model’s Own Rollouts
Abstract
Interpretability work on world models reads latents with probes calibrated on encoder outputs, then reuses them on latents the model generates with its own dynamics; read that way, the intervention effects a model imagines have been found not to match the environment’s. We ask whether that is lost information or a change in the measurement’s own calibration, and whether that calibration is stable across rollout conditions, holding a TD-MPC2 model fixed and varying only how the probe is calibrated. The measured physical variables stay linearly recoverable: on perturbed cheetah states, a map from rolled to encoded latents fit without the targets restores an encoder-calibrated readout from to . The map is specific to the condition it was fit in: it keeps only part of its benefit at other horizons, can do worse than no map on another state distribution, and is poorly captured by an orthogonal map, while identity plus a rank-16 correction recovers most of its gain. Recalibration also changes intervention-consistency effect sizes, by up to across four comparisons, without reversing any feature-versus-control separation; in the comparison a post-hoc check finds both readouts above on the corresponding states. In a second architecture, LeWorldModel on Reacher, the encoder-calibrated readout stays almost perfectly accurate on the model’s training distribution, so the large TD-MPC2 mismatch is not universal across the settings tested.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.