acceptodds
Under review as a conference paper at ICLR 2027

Unsupervised Generative Replay for Efficient Reinforcement Learning

Abstract

Leveraging the diffusion model to augment the replay buffer has become a useful paradigm to enhance the efficiency of reinforcement learning (RL). Can we pre-train a diffusion model from unlabeled exploratory data such that it can be adapted to new downstream tasks to enrich the RL buffer, even when the downstream reward function cannot be queried on these pre-collected data? In this work, we present an Unsupervised Generative Replay (UGR), a reusable generative replay framework for this setting. UGR pre-trains a diffusion model on exploratory transitions relabeled with diverse reward priors to serve as unsupervised generative priors. When adapting to a specific task, we only train a lightweight guidance model to steer data generation towards the target task. Experiments across complex control tasks on DeepMind Control Suite reveal that UGR effectively improves online data efficiency under a fixed interaction budget. These results suggest that reusing unsupervised generative priors is a promising way to improve online RL across downstream tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.