acceptodds
Under review as a conference paper at ICLR 2027

Resolution Collapse in Latent-Space CEM Planning

Abstract

Planning with a learned latent world model usually means using the cross-entropy method (CEM) to search over candidate action sequences and ranking them by how close the predicted final latent state lands to the goal. We show that this ranking step fails in a specific and measurable way. Across four continuous-control environments built on a joint-embedding predictive world model, the correlation between predicted and true candidate quality collapses to chance exactly at the terminal frame CEM optimizes, even though the same predictor ranks candidates well several steps earlier. We call this resolution collapse, and we observe it on every environment we test. A set of horizon-manipulation experiments then shows that the collapse tracks CEM's own optimization target rather than a fixed prediction horizon. Stretching the planning horizon from five steps to eight steps moves the collapse point outward by exactly the same amount, and the correlation at whichever frame is currently terminal stays pinned near chance at every horizon we try. This motivates a direct fix. Ranking candidates on a frame the model can still resolve, instead of the terminal one, raises planner success rate by three to fifteen points across our environments, following an inverted-U pattern in which asking too early actively hurts and asking too late is where the original failure lives. Two independent random-cost controls, matched in magnitude to the real cost, rule out that this gain is a side effect of simply adding another term to the objective. Replacing the resolution-matched cost with noise of identical spread collapses success rate by twenty-three to ninety-four points, and an additive version of the same idea beats a matched-noise version by eleven to thirteen points at p<0.001. Finally, we give a rule that selects the reporting frame directly from the model's own resolution curve, without ever consulting success rate, and show it lands within two points of the oracle frame found by an exhaustive sweep, on every environment. Together, these results turn an invisible bottleneck in latent-space planning into something we can measure, predict, and correct at no extra test-time cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.