Nested Latent Representations for Size-Flexible End-to-End World Models
Abstract
Latent world models for planning typically assume a fixed representation size. While recent work, such as SubJEPA, considered operating in subspaces of smaller dimensionality, they require additional task-specific tuning of the subspace granularity. In this work, we ask whether world models can learn representations that remains useful across multiple post-hoc embedding sizes. As an answer, we introduce ElasticJEPA which jointly trains a world model at multiple nested granularities. ElasticJEPA combines subspace regularization with prediction losses applied to the same prefixes used at inference. The novel prediction loss objective induces a coordinate-supervision profile that is naturally monotonically-ordered. ElasticJEPA also ensures that privileging prefixes are retained under truncation. With such a construction, a single trained model can be evaluated at multiple representation sizes with no need for retraining. Importantly, inference-time evaluation can also include widths not supervised during training. Across four environments, ElasticJEPA is competitive with prior methods at full width, retains most of its performance under substantial truncation, and degrades less than baselines truncated to the same size, most clearly at the smallest widths.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.