SAFE-HAC: A Modular Architecture for Safe Reinforcement Learning
Abstract
Safety information is heterogeneous across continuous-control tasks: some environments expose only a cost signal, others provide signed constraint margins or task-specific regulated outputs, and only some admit a usable conventional controller. Existing methods are typically specified around one such interface. We study safe policy learning under optional safety signals and optional controller assistance. We introduce SAFE-HAC (Soft Actor with Fallback Enhancement and Horizon-Aware Critics), a modular recurrent off-policy actor–critic whose availability mask activates any subset of three finite-horizon predictors: cumulative cost, worst future clearance, and regulated-output error. The last signal can represent distance, speed, attitude, or another task-defined objective; it is Lyapunov-inspired only when its zero set and descent condition have the corresponding control-theoretic meaning. Active predictions enter the actor loss. When a fallback controller is available, the same predictions assess both proposed and fallback actions and can reject a policy proposal in favor of an MPC, LQR, PID, or other conventional control action; without a fallback, learning remains well-defined and the policy executes directly. Because intervention changes the transition-generating action, critics train on the executed action while proposal–fallback pairs support intervention imitation. We formalize the error induced by proposal-labeled shielded replay. We evaluate both policy-only and controller-assisted modes on five continuous-control tasks against SAC-Lag, PPO-Lag, CPO, and TD3. Empirical evaluations demonstrate that SAFE-HAC achieves competitive task performance while reducing constraint violations and reaching the specified empirical safety criteria in fewer environment steps compared with standard constrained-RL baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.