Reweight, Do Not Re-calibrate: Reusing the Same Encoder-Predictor Pair (World Model) When the Policy or Query Changes
Abstract
A world model calibrated under one policy is queried under another once a planner selects actions from its predictions. We ask whether the same frozen encoder-predictor can be reused under a changed policy or a changed query law without re-calibration on target labels. Conformal off-policy prediction marginalizes the action and predicts a return, which makes the importance weight a function of the outcome and forces an outcome model, rejection subsampling, or a trajectory product. We condition on the action and predict the next state, which makes the weight measurable before the outcome and removes all three; the reduction holds at one step, and the occupancy factor returns over a horizon. With a planner enacting the new policy on RoboCasa, exact propensity weighting of the source calibration set reaches 0.913 coverage with no new observations, against 0.910 after recalibrating on 75 labelled target episodes per task (model seed 271828, temperature 3); recalibration re-collects 300 labelled episodes under each new policy, while the weighted construction reuses one 600-episode source set across all nine target policies. We prove that the target-to-source likelihood ratio is measurable before the outcome and factors into state-occupancy and action-propensity ratios, and combine weighted conformal prediction with the exact or an estimated ratio. Under a query change, coverage at nominal 0.90 moves from 0.902 to 0.550 and to 0.909 with learned weights on 192 Franka transitions, and from 0.734 to 0.903 on 5,100 FurnitureBench episodes, matching oracle weighting; with V-JEPA-2-AC in place of DINOv2 on the Franka data, 0.561 and 0.917. Under the planner-enacted policy change, coverage moves from 0.750 to 0.912 with each radius finite, a paired gain of 16.2 points [8.5, 23.7], replicating on two further model seeds at 0.690 to 0.877 and 0.735 to 0.892. For a learned ratio, the exact convex dual of the classifier-regret constraint within each quantile bin of the posterior gives coverage lower bounds of 0.863 and 0.802 on the two query-change corpora, against observed 0.909 and 0.903. Finite-support identities of the argument are machine-checked in Lean 4.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.