acceptodds
Under review as a conference paper at ICLR 2027

Supervise What Decides Success: Criterion-Aligned Auxiliary Losses for Latent World-Model Planning

Abstract

Latent world models plan by scoring candidate action sequences with distances in latent space. However, task success is judged by physical quantities, hereafter the success-criterion quantities. In all four latent world models we examine, the end-effector position is encoded in the latent state with an error larger than the success criterion allows. Such a latent state cannot separate successful candidates from failing ones. We propose an auxiliary loss that uses success-criterion quantities as training targets, whereas existing latent world models take them only as inputs. During training, a linear head on the encoder and predictor outputs regresses the success-criterion quantities, and the regression error is added to the training loss. The head is discarded after training, so the model, its cost, and its inputs at inference are unchanged. This loss alone improves the success rate on PushT and cube by 3.5% and 3.4% (absolute), and both improvements are statistically significant. A success criterion thus specifies what a world model must retain in its latent state, and we show that it can serve directly as a training target.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.