acceptodds
Under review as a conference paper at ICLR 2027

Flow-Lag: Sensitivity-Aware Primal-Dual Optimization for Safe Flow Policies

Abstract

Flow-based policies can represent expressive action distributions and have recently shown strong performance in reinforcement learning. However, enforcing safety constraints for flow-based policies remains challenging. Most flow-policy optimization methods optimize surrogate objectives defined over the internal flow process, such as flow-matching or path-space objectives, rather than the deployed endpoint policy directly. This creates a surrogate-endpoint mismatch in which the policy is updated in surrogate objectives, while safety is ultimately determined by the endpoint actions executed in the environment. We show that this mismatch can substantially weaken safety correction in primal-dual reinforcement learning. Specifically, increasing the Lagrange multiplier can induce substantial changes in the flow policy while yielding only limited reductions in cost. To address this issue, we propose Flow-Lag, a sensitivity-aware primal-dual framework that propagates endpoint cost information through the flow dynamics to measure the safety response of the actual policy update. Flow-Lag uses this response to adapt the multiplier and calibrate the policy update, directly linking primal-dual correction to its endpoint effect. Across Safety-Gymnasium navigation and velocity-constrained Safe MuJoCo benchmarks, Flow-Lag achieves improved reward-safety performance over representative flow-based and conventional safe RL baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.