Plan Toward Success, Not Similarity: Success-Aligned Costs on Frozen Latent World Models
Abstract
A frozen latent world model serves control through test-time search, where an optimizer seeks the plan of lowest cost. We argue that the cost is the main lever, and that latent and temporal distances measure the wrong quantity, embedding similarity or elapsed steps. We propose success-aligned distances, heads that regress the logged time until the task predicate first holds, a label that is zero exactly on successful states. SAD-CEM optimizes a pessimistic ensemble of these heads toward the encoded goal with cross-entropy search, and consults generated waypoints when its plan leaves the generated paths. With frozen LeWM backbones, SAD-CEM has the highest or tied-highest mean success on each of five goal-reaching environments, and a five-environment mean of 0.794 against 0.630 for the strongest of fifteen external planners. The success-aligned cost alone, without generated waypoints, already reaches 0.742, while a step-gap label or one head in place of the ensemble lowers success significantly. Code and demos are available at anonymous.4open.science/r/sad-cem.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.