acceptodds
Under review as a conference paper at ICLR 2027

From Representation Learning to Reshaping Latent Geometry in World Models

Abstract

Latent world models encode high-dimensional observations into compact representations, where the dynamics can be learned and used for planning. Their representation quality is typically evaluated based on few-step prediction accuracy, while goal-conditioned planning aims to minimize the distance between long-term predictions and an encoded goal. In this work, we show that without proper supervision, the representation's *metric consistency* is poor, meaning that latent distances do not align well with the optimal cost of reaching the goal. We then reshape latent geometry using these costs and generalize the method through temporal ordering within recorded trajectories. Next, we show that geometry reshaping does not require extensive visual representation learning. We introduce an inexpensive test for representation sufficiency, based on the recovery of task-relevant state variables. We show empirically that once this criterion is met, visual training can stop and geometry reshaping in latent space alone can significantly improve planner performance. After two epochs of visual training in TwoRoom, the representation already satisfies our sufficiency criterion, but CEM planning success is only 11%. Freezing the encoder and continuing training in latent space raises success to 88% with proportional task-distance supervision and 90% with temporal ordering, compared with 27% using the base objective. Continued expensive visual training without reshaping plateaus at 37%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.