DRIVE: Decoupled Representation with Informed Value Estimation
Abstract
Model-based reinforcement learning (MBRL) promises sample-efficient continuous control, yet existing methods face a fundamental tension between representation quality and training stability. Reconstruction-based approaches allocate encoder capacity to value-irrelevant features, while joint dynamics-value supervision lets noisy bootstrapped gradients distort the learned latent geometry. We introduce DRIVE (Decoupled Representation with Informed Value Estimation), a framework built on one principle: representation learning should be decoupled from value-driven optimization. DRIVE isolates the encoder from all critic and actor gradients, training it exclusively through structured supervision, including dynamics prediction, contrastive temporal consistency, reward estimation, and cumulative return prediction on the unit hypersphere. The critic fits values on the resulting well-conditioned latent space, and the actor combines single-step Q-gradients with a multi-step soft-decay rollout through the learned dynamics, internalizing trajectory-level planning at training time with no additional inference cost. Across 28 high-dimensional continuous control tasks, DRIVE matches or exceeds state-of-the-art model-based and model-free baselines. Ablation studies confirm that the soft-decay rollout and structured encoder supervision provide complementary gains. Crucially, our analysis reveals a fundamental failure mode in high-dimensional MBRL: local representation aliasing inevitably leads to Q-value flattening and actor gradient starvation, a pathology that DRIVE effectively cures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.