acceptodds
Under review as a conference paper at ICLR 2027

Asymmetric Actor-Critic Pretraining for Efficient Flow-Based Online RL Fine-tuning

Abstract

Offline-to-online (O2O) reinforcement learning employs synchronized, symmetric offline training of the actor and critic for downstream online fine-tuning. Our analysis reveals that this symmetric training process impairs performance during online fine-tuning: while the policy network acquires a sampling prior through offline training to ensure efficient online fine-tuning, the value network may lose plasticity due to prolonged offline training, thereby failing to rapidly adapt to new environments when faced with distribution shifts between the offline and online settings. Based on this distinction, we identify a trade-off mechanism regarding critic offline training and investigate the impact of offline data coverage on the actor's online performance. Our research reveals an asymmetry in offline training: the critic is highly sensitive to the extent of offline training, whereas the actor relies more heavily on the coverage of effective regions within the offline data. Leveraging this asymmetry, we propose Critic-Free Pretraining—a simple, asymmetric training framework widely applicable to various algorithms. CFP involves pretraining the actor without a critic, subsequently initializing a fresh critic with a joint offline warm-up before entering the online interaction phase. Across 8 challenging manipulation domains, spanning OGBench and Robomimic, CFP achieves performance comparable to or superior to naive O2O methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.