Can In-Sample Offline RL Learn Critics First and Extract Policies Later?
Abstract
Recent progress in robotics and foundation models is driving offline reinforcement learning toward increasingly expressive but computationally expensive policy classes, including diffusion-based, flow-based, and large VLA policies. In-sample value-learning methods such as IQL and XQL are appealing in this setting because they estimate value functions without sampling actions from the current policy. This suggests a simple two-stage strategy: first perform value learning, then freeze the learned critic and extract a policy afterward, which we refer to as the *Freeze-then-Extract* protocol. However, existing in-sample value-learning methods typically still rely on alternating policy and critic updates in practice. Motivated by this gap, we systematically evaluate Freeze-then-Extract across offline RL methods and benchmarks and find that, despite its appeal, it often underperforms standard alternating updates. To understand this degradation, we investigate how the learned critic evolves throughout training. We find that advantage estimates vary substantially across training steps, even toward the end of critic training. When policy extraction relies on a single frozen critic, these fluctuations can lead to transient high-advantage peaks and a reduction in the effective number of training samples. We therefore investigate whether aggregating signals from multiple frozen critics, obtained from different checkpoints or independent initializations, can stabilize policy extraction while preserving policy-sampling-free value estimation. We find that such ensembles improve average extraction performance. In VLA experiments, critic-first training reaches strong performance with fewer expensive policy updates. Our results identify temporal variability in advantage estimates as a key obstacle to reusable frozen critics in scalable offline RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.