LeDreamer: Bridging the Understanding of Representation Collapse Between JEPA and Sample-Efficient World Models
Abstract
Joint Embedding Predictive Architectures (JEPAs) efficiently learn representations without pixel-level reconstruction, with recent methods seeking to replace stabilization heuristics such as stop-gradient with the minimization of well-defined objectives. Although the Sketched Isotropic Gaussian Regularizer (SIGReg) prevents representation collapse in continuous JEPAs, its direct application does not account for the factorized structure of discrete categorical latents used by sample-efficient architectures such as DreamerV3. In these latents, shifts applied to the logits of each categorical distribution do not change the resulting probabilities, but can still affect the regularization objective. We introduce LeDreamer, a decoder-free transformer-based world model that addresses this mismatch by applying SIGReg per distribution to pre-sampling categorical logits. We prove that applying SIGReg to the flattened representation, formed by concatenating the logits of all categorical distributions, admits zero-information solutions whose regularization cost decreases as with the number of categorical distributions , whereas applying SIGReg independently to each distribution assigns every such solution a positive cost independent of . Empirically, per-distribution SIGReg prevents marginal representation collapse where flattened regularization fails. We further find that preventing marginal collapse is insufficient for control: in this setting, retaining stop-gradient remains necessary to learn sufficiently action-sensitive dynamics and achieve competitive agent returns. To systematically evaluate these effects, we introduce Atari-5-100k, a statistically selected five-game ablation suite derived from results for 32 recent sample-efficient world models. Across the complete 26-game Atari 100k benchmark, LeDreamer recovers 79.7% of the interquartile mean score of DreamerV3 while training 2.12 times faster, performing on par with recent decoder-free alternatives while training between 2.0 and 3.8 times faster than evaluated decoder-enabled architectures.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.