PRISM: Probing Physical Consequence Signals In World Model Predictions
Abstract
Action-conditioned world models predict the consequence of each candidate action, yet standard planners reduce each prediction to a single feature distance to a goal image. This scalar readout couples task progress with obstacle safety in one number, and the progress signal dominates: consequence information that the model faithfully predicts is lost in the aggregation. We expose this readout bottleneck through a controlled comparison. A frozen text-image projection, aligned once on robot data without collision labels, reads named object relations from the same predictions. Where feature-distance planning fails on obstacle-dense pick-and-place sequences, this relational readout achieves 71.7%, revealing that the learned representations already encode physical consequences. Three isolating interventions confirm the information is per-candidate and prediction-dependent: shuffling action-score pairings eliminates every gain, reading from current observations instead of predictions yields only partial benefit, and goal-image contrast rules out directional matching. This is a representation-level finding: the consequence information exists regardless of how downstream planning uses it. PRISM exploits this by decoupling direction from selection: geometry handles progress toward the goal, and two sequential consequence gates read VLM-written relations from each candidate’s predicted scene, filtering harmful candidates before geometry picks among survivors. On 450 transport trials, consequence gating raises completion from 62.2% to 86.7% and cuts obstacle displacement by 63%, lifting safe completion from 33.3% to 55.6%. The mechanism transfers to pick-and-place sequences and to a real Franka arm without retraining
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.