Sample-Anchored Advantage Weighting for Flow Policy Learning
Abstract
Offline-to-online reinforcement learning combines offline pretraining with online fine-tuning, resulting in replay buffers that mix offline data from diverse behavior policies with online experience from the evolving policy. Flow policies are well suited to capturing multimodal action distributions through iterative generation, but their implicit action distributions make conventional RL objectives difficult to apply, leaving behavior cloning as a natural yet limited training strategy. To improve performance, advantage-weighted regression (AWR) retains the imitation objective while weighting it according to estimated advantages. However, this weighting compares actions across modes on a shared value scale, which can suppress valid behavior modes whose estimated values are temporarily lower and thereby reduce behavioral diversity. To address this issue, we propose Sample-anchored Advantage-Weighted Regression (SAWR), a variant of AWR for flow policies that replaces conventional global value comparison with a sample-specific local comparison against a reference action constructed from each replay action. Specifically, SAWR constructs the reference action by partially noising the replay action and completing it with the flow policy, then weights the flow matching loss by the value difference between the replay action and its reference. This sample-specific comparison uses local value differences to guide policy learning while mitigating the suppression of valid behavior modes. Integrated into Flow Q-Learning, SAWR improves average online success rates over conventional AWR on OGBench and D4RL Adroit. Navigation and manipulation diagnostics further demonstrate better preservation of diverse successful behaviors, while a route-blockage test shows improved robustness relative to conventional AWR when previously preferred behaviors become unavailable. Code is provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.