Beyond Safe Demonstrations: Failure-Informed Flow Policies for Offline Safe Reinforcement Learning
Abstract
Can an offline learner satisfy a deployment cost budget without complete trajectories whose cumulative cost is within that budget? Bootstrapped cost estimates may fail to capture delayed costs that only become visible in recorded subtrajectories. We introduce FIBA-Flow, a failure-informed, budget-adaptive approach to offline safe reinforcement learning. For each candidate action, we combine a cost-to-go estimate with a short-horizon cumulative-cost prediction to obtain a cost score. We optimize target action probabilities to maximize entropy-regularized predicted return under a remaining-budget constraint on the mean cost score. A budget-conditioned flow is trained to match these target distributions for one-step action generation. A guide learns to adjust the flow's internal budget input through action-distribution matching. Its targets are constructed across budget conditions under a constraint on the predicted cost set by the actual remaining budget, using critics that account for subsequent guide decisions. Evaluations on 38 DSRL tasks show that FIBA-Flow satisfies the cost constraint on every task and achieves the highest normalized return on 25 tasks among budget-compliant methods. Our anonymous link is: https://anonymous.4open.science/r/paper-6d0be256/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.