acceptodds
Under review as a conference paper at ICLR 2027

Offline Quality-Diversity

Abstract

Offline reinforcement learning methods learn a single policy from a static dataset, but in doing so fail to leverage the behavioral diversity the dataset may contain. Meanwhile, Quality-Diversity (QD) methods produce rich archives of diverse high-performing behaviors, but require online environment interactions, which poses a fundamental barrier when trial-and-error is costly or unsafe. In this paper, we try to reach the best of both worlds: the same dataset that yields a single offline RL policy could, in principle, yield an entire repertoire that QD methods could find. We introduce Offline Quality-Diversity (Offline QD), a new problem setting in which the goal is to recover a full behavioral repertoire over a given descriptor space from a static dataset, without any environment interaction during training. We propose three independent methods that each solve this differently: successor features that act as a proxy for a policy's achieved descriptor; a learned world model that lets a standard QD algorithm search entirely in imagined rollouts; and a behavior-conditioned decision transformer that recovers an archive through sequence modeling alone. Experiments show that Offline QD recovers a substantial fraction of online QD performance on tasks where exploration is easy, and substantially surpasses online QD on tasks where random exploration is hard but the offline dataset provides more diversity. We further show that offline archives serve as a strong initialization for online search, turning environments where online QD struggles from scratch into tractable ones. Our results suggest that offline datasets already contain far more behavioral diversity than standard offline RL methods extract, opening a new direction for QD research beyond online exploration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.