acceptodds
Under review as a conference paper at ICLR 2027

UniQ: Unified Conditional Flows for Offline Safe Reinforcement Learning

Abstract

Offline safe reinforcement learning seeks to maximize reward under cost constraints using a fixed dataset. Lagrangian methods repeatedly adjust a multiplier, changing the policy objective and potentially destabilizing learning when cost estimates are inaccurate. We propose , a conditional flow framework that replaces continued multiplier adaptation with a fixed safety rule estimated from offline data. The rule prioritizes reward in safe regions, balances reward and cost near the safety boundary, and prioritizes cost reduction in unsafe regions. Conditioned on state and accumulated cost, a single flow critic learns the combined return and guides a conditional flow policy, with both models conditioned on the physical state and accumulated cost. To train these models across a wider range of budget conditions, we resample cumulative costs for each logged transition. Our analysis establishes population Bellman consistency and fixed-policy convergence under vanishing update error, and identifies the utility-change term eliminated by the fixed safety rule. A capacity-restricted construction further demonstrates a strict value-error advantage over separately fitted component critics. On nine MetaDrive tasks, satisfies the cost threshold on every task and achieves approximately higher average normalized reward than the strongest baseline. [Anonymous Project Page & Videos](https://anonymous4699-uniq-project.static.hf.space/)

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.