acceptodds
Under review as a conference paper at ICLR 2027

Sparse-TD-JEPA: Improving Zero-Shot Reinforcement Learning With Sparse Latent Task Representations

Abstract

Successor-feature-based zero-shot reinforcement learning methods pre-train a task encoder and successor-feature predictor on offline, reward-free data, then recover a policy at test time by solving a closed-form linear regression over a small labeled dataset. We find that the success rate of this method can be improved by a property that has received little attention: when the task encoder is dense, the predicted reward surface is less concentrated near the goal, producing a weaker gradient for the policy to follow. We address this by introducing a sparsity learning objective into the task encoder that concentrates reward-predictive structure into a smaller effective subspace of the task encoder's latent space. The sparsity penalty is applied during training and requires no change to the closed-form inference procedure at test time. We instantiate this idea as Sparse-TD-JEPA, a direct extension of the recently published TD-JEPA framework, a Temporal Difference (TD) reinforcement framework for the Joint-Embedding Predictive Architecture (JEPA). We evaluate our method across various standard zero-shot performance benchmarks covering various robot manipulation, locomotion, and exploratory navigation tasks. Our method does not only achieve consistent better zero-shot performance than the baseline framework, it achieved the best performance of all established zero-shot baselines tested in most cases.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.