VE-JEPA: Joint Latent and Value Learning for Offline Reinforcement Learning
Abstract
Predictive representations are typically assessed by how accurately they predict. We examine whether predictive accuracy is informative about downstream control in offline reinforcement learning. We introduce VE-JEPA, a planning-free joint-embedding predictive architecture in which a single encoder is trained through latent prediction, reward prediction, and expectile value learning, with an advantage-weighted policy operating directly on encoded observations. Across three medium-replay locomotion tasks and five seeds per condition, allowing Q and V losses to update the encoder substantially improves control over prediction-only and prediction-plus-reward encoder training. Ablations show that latent prediction improves reward prediction from frozen encoder features, while its incremental control benefit varies across tasks and seeds. Encoder learning rate also affects control performance. Reducing the encoder learning rate tenfold raises mean AntMaze medium-play success from 5.0% to 36.0%, averaged across seven training seeds. Measurements on fixed observations show reduced representation movement despite larger encoder-gradient norms and higher predictive loss. Extensions across locomotion datasets reveal a training-budget tradeoff: slower encoder updates can delay performance improvements at a fixed budget. Across nine D4RL locomotion datasets with ten seeds per method and dataset, VE-JEPA achieves an aggregate interquartile mean normalized return of 74.7 against Decision ConvFormer's 71.5 at 200,000 training steps under fixed-target evaluation, a difference of with a 95% bootstrap confidence interval of . This comparison varies across data regimes: VE-JEPA leads by 13.8 points on medium-replay and trails by 8.8 on medium-expert; the overall ordering is sensitive to target-return selection. These results highlight the importance of value-informed encoder training and task-dependent encoder adaptation when learning predictive representations for offline control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.