acceptodds
Under review as a conference paper at ICLR 2027

C-Steering: Constrained Noise-Space Steering of Behaviour-Cloned Flow Policies for Offline Safe RL

Abstract

Offline safe reinforcement learning must find high-return behaviour within an episodic cost budget from logged data alone. We propose C-Steering, a bounded noise actor over a one-step flow decoder trained only by behaviour cloning. The actor selects actions from the learned decoder's image without a competing cloning penalty. Controlled experiments examine three design choices: cost-budget calibration, decoder-prior geometry and noise-actor parameterisation. Calibrating discounted cost values to the episodic budget improves budget satisfaction substantially more than the tested multiplier-update changes. A bounded full-dimensional prior raises mean return by over a covariance-matched sphere while retaining of budget-satisfying cells, with gains in both return and seed-level safe return supported by seed and task resampling. Additional diagnostics reveal near-boundary noise selection and an approximately larger RMS spread of the decoded boundary image on two-dimensional tasks, characterising how the learned action geometry differs between priors. In a single-task comparison, a direct actor achieves higher return at every seed and lower mean cost than a reflected-flow actor, with two network passes per action instead of eleven. Across DSRL tasks and two budgets, C-Steering meets the budget on cells, with positive return, compared with and , respectively, for the baseline with the most budget-satisfying cells. These findings establish an effective constrained-steering design and provide empirical insight into its prior geometry and noise selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.