Episodic Control for Safe RL with Guarantees
Abstract
On-policy safe reinforcement learning (RL) methods provide strong stability and constraint guarantees but are sample-inefficient, while off-policy methods improve data efficiency at the cost of stronger assumptions to obtain guarantees, creating a fundamental tension in safety-critical settings. We introduce the Safe Integration Framework (SIF), a class of admissible critic corrections that inject bias into on-policy value estimation while preserving contraction of the Bellman operator and bounding deviation from the nominal fixed point. We further show that SIF preserves constrained trust-region optimization guarantees up to an explicit slack term, and yields primal-dual convergence to a bounded neighborhood of the nominal saddle point. As an instantiation, we propose Safe Proximity-guided Episodic Control (SPEC), a lightweight retrieval-based plugin that augments on-policy cost critics with uncertainty-aware corrections. Across six maximum-difficulty Safety-Gymnasium navigation tasks spanning the Button, Goal, and Circle families, SPEC applied to TRPO-PID yields a more consistent constraint profile than all baselines while avoiding reward collapse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.