acceptodds
Under review as a conference paper at ICLR 2027

A Single Occupancy Measure Is Often Enough for Zero-Shot RL

Abstract

As unsupervised pretraining becomes increasingly ubiquitous in reinforcement learning, developing a theoretical understanding of these methods is becoming as important as their empirical success. We focus on unsupervised learning via interaction, where forward-backward (FB) representation learning serves as a prototypical example. In standard FB, the policy defining the successor measure depends on the representation being learned, resulting in a coupled system. This raises the question: *Can the successor measure of a single behaviour policy suffice for zero-shot RL?* To understand the tradeoff this simplification introduces, we build on prior work that analyzes FB and the finite-dimensional factorizations of successor measures. We derive additional necessary conditions for exact FB representations, which we use to study learning behaviour in didactic settings, and quantify the consequences of low-rank approximation used in practice. In didactic settings, we observed substantially improved convergence when learning with smaller representation errors than FB. However, in a didactic counterexample, we observed that even perfect knowledge of need not yield optimal downstream control. We then compare our simplified method, One-Step FB, to prior work across state-based and image-based continuous control domains. One-Step FB yields representations achieving the zero-shot performance of FB on average. We further show that the zero-shot policies inferred by our algorithm provide an efficient initialization for further fine-tuning on downstream tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.