acceptodds
Under review as a conference paper at ICLR 2027

DT-JEPA: Learning to Predict Future Representations for Decision-Making

Abstract

Decision Transformer (DT) has demonstrated strong performance in various offline reinforcement learning (RL) tasks by conditioning action generation on expected future returns and historical trajectories. However, treating expected future returns as the condition provides an incomplete description of the future, as trajectories with different state transitions and behaviors may share similar returns. Moreover, expected future returns inherently depend on task-specific reward functions, which makes the policy learning reliant on reward-labeled datasets and restricts its scalability to tasks with similar but unseen rewards. Recently, Joint Embedding Predictive Architectures (JEPA) has emerged as a promising paradigm of self-supervised learning-based world models, which learns predictive representations of future dynamics in latent space rather than reconstructing future states in the original space. Inspired by JEPA, we propose DT-JEPA, an offline pretraining and online finetuning framework that learns general knowledge from reward-free offline trajectories and achieve excellent adaptation to the unseen tasks. Instead of predicting a single deterministic future embedding, we design a novel probabilistic predictor to model the conditional distribution of future representations in the latent space, thereby capturing both predictive information and uncertainty about future dynamics. The predicted future representation is then utilized to guide DT's action generation. To generalize to tasks with unseen reward objectives, we further develop a critic-guided fine-tuning mechanism to efficiently adapt the pretrained policy to unseen reward functions with limited online interactions. Theoretically, we show that 1) the JEPA objective seeks to maximize a variational lower bound on the mutual information between historical representations and future trajectories; and 2) predictive future-conditioned policy learning admits a finite-sample performance bound with an estimation term. Experiments on D4RL benchmarks demonstrate that DT-JEPA achieves efficient adaptation across varying offline RL tasks, outperforming traditional DT-based and other offline RL methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.