Asymmetric Actor-Critic Pretraining for Efficient Flow-Based Online RL Fine-tuning
Abstract
Offline-to-online (O2O) reinforcement learning employs synchronized, symmetric offline training of the actor and critic for downstream online fine-tuning. Our analysis reveals that this symmetric training process impairs performance during online fine-tuning: while the policy network acquires a sampling prior through offline training to ensure efficient online fine-tuning, the value network may lose plasticity due to prolonged offline training, thereby failing to rapidly adapt to new environments when faced with distribution shifts between the offline and online settings. Based on this distinction, we identify a trade-off mechanism regarding critic offline training and investigate the impact of offline data coverage on the actor's online performance. Our research reveals an asymmetry in offline training: the critic is highly sensitive to the extent of offline training, whereas the actor relies more heavily on the coverage of effective regions within the offline data. Leveraging this asymmetry, we propose Critic-Free Pretraining—a simple, asymmetric training framework widely applicable to various algorithms. CFP involves pretraining the actor without a critic, subsequently initializing a fresh critic with a joint offline warm-up before entering the online interaction phase. Across 8 challenging manipulation domains, spanning OGBench and Robomimic, CFP achieves performance comparable to or superior to naive O2O methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.