PACE: Path-Aware Constraints on Exploration in Frozen Flows for Offline-to-Online Reinforcement Learning
Abstract
Flow policies capture complex, multimodal offline behaviors, yet direct value optimization through iterative generation can destabilize training. Existing methods instead train a flow teacher to imitate offline data and guide the actor that interacts with the environment. During the online phase, however, updating this teacher on newly collected experience can undermine its learned behavioral structure and destabilize policy improvement. Preserving a stable behavioral reference while continually improving its guidance remains an open problem. To address this, we introduce Path-Aware Constraints on Exploration (PACE) in frozen flows for offline-to-online RL, which freezes the pretrained flow and uses an evolving critic to guide exploration in its latent noise space. The explored and original outputs jointly supervise the actor, enabling a frozen teacher to offer both stable behavioral guidance and continually updated knowledge. Using the flow's fixed geometry, PACE further constrains deformation along complete noise-to-action paths to reduce out-of-distribution risk. Experiments on 35 OGBench tasks demonstrate improved sample efficiency and final performance under limited online interaction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.