Covering Behaviors, Not Latents: D-Optimal Curricula for Zero-Shot Reinforcement Learning
Abstract
Zero-shot reinforcement learning aims to learn a reusable policy family before downstream rewards are specified. Recent methods have advanced this goal through predictive representations, which delegate the coverage of behaviors to randomly sampled latent variables or task directions. However, uniform sampling in this index space does not guarantee uniform coverage of the behaviors that matter for transfer. We study this gap through the coverage of policy embeddings, defined by the discounted feature expectations induced by each policy. Motivated by an analytical link between coverage geometry and zero-shot suboptimality, we propose a D-optimal coverage criterion that favors policy embeddings spanning a well-conditioned region of feature space. While leaving each policy’s task-aligned objective unchanged, the corresponding leverage score defines a training curriculum centered on expanding under-covered directions in embedding space. Empirically, we show that D-optimal coverage tracks zero-shot performance and that the resulting curriculum improves learning with both exact and learned policy embeddings, including offline transfer on ExORL, where the gains are most consistent under limited data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.