SAFTER: Safety-Aware Filtered Training for Efficient and Robust Reinforcement Learning with Closed-Form Backup CBFs
Abstract
Safe reinforcement learning under bounded actuation requires more than preventing an immediate constraint violation: a currently admissible action can still leave insufficient control authority for future recovery. Backup control barrier functions (bCBFs) encode this finite-horizon recoverability, but sampled backup prediction introduces many simultaneous current, future, terminal, and input-admissibility conditions. Repeatedly solving the resulting optimization inside RL training is undesirable, and an opaque or approximated projection Jacobian can also give the actor a mismatched view of how its actions are actually executed. We propose SAFTER (Safety-Aware Filtered Training for Efficient and Robust reinforcement learning), which converts sampled backup recoverability into a differentiable analytical training map. SAFTER augments the applied command with first-order dynamics, composes current-state, backup-horizon, terminal, and continuous-model input-admissibility conditions with a hierarchical Log-Sum-Exp certificate, and reduces the resulting CBF condition to a single affine half-space with an exact minimum-intervention projection. In the sample efficiency study, SAFTER reaches 90% training success after 4.06M transitions, versus 5.85M for Vanilla SAC and 6.16–6.26M for conventional CBF-filtered SAC, while closely matching the 4.13M transition count of bCBF–LSE–QP; this indicates predictive recoverability, rather than the closed-form realization alone, is primarily associated with the sample-efficiency gain. Under Low-, Medium-, and High-Shift geometries, SAFTER achieves 98.8%, 88.6%, and 61.4% success, improving over Vanilla SAC by up to 67.8 percentage points and over CBF–CF by up to 40.2 points. With the same correction reward, replacing the exact projection Jacobian by the identity reduces High-Shift success from 61.4% to 6.2%. These results show that backup recoverability improves learning efficiency, while the analytical, differentiable realization is important for learning policies that remain effective under geometric distribution shift.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.