What Data Is Needed for Offline RL in Continuous Control? A Geometric Perspective
Abstract
Learning from high-quality data forms the foundation of most modern AI applications, ranging from large language models to generalist robotic systems. However, collecting sufficient expert data to satisfy the increasing demands of robust performance is infeasible in general. In this paper, we study the question: what properties must offline data exhibit to enable near-optimal behavior in continuous control? We present a rigorous theoretical analysis of the offline RL problem, and outline the inability of common notions of coverage to discriminate between benign and intractable problem instances. To this end, we propose new notions of coverage which capture the inherent geometric nature of continuous control. Via novel analyses leveraging stability-theoretic tools and these geometric notions of coverage, we identify regimes where offline RL is tractable versus difficult, and we demonstrate how popular model-based algorithmic and data practices can achieve polynomial regret bounds under these regimes. We perform experiments on manipulation and locomotion tasks and find strong alignment with our theoretical predictions, as well as empirical evidence that our recommended practices benefit a wider range of model-based and model-free frameworks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.