Verification-Guided Safety-Repair of Learned Controllers
Abstract
Formal verification can certify that a learned controller is safe. But what should be done when verification fails? Simply suppressing actions that lead to unsafe next states is insufficient because the dangerous action that ultimately caused the safety violation might have occurred many steps earlier. Retraining modifies the policy globally, thereby potentially introducing new failures. Instead, we introduce ReachECO, a verification-guided repair procedure that exploits the information of failed multi-step reachability analysis for synthesizing explicit constraints on controller behavior. For Differentiable Weightless Controllers, the repairs compile into exact circuit patches that leave the controller unchanged everywhere outside the targeted policy cells. We prove soundness and termination of the repair procedure for ReachECO, as well as, assuming exhaustive continuation search, completeness relative to the reachability abstraction. Empirically, we demonstrate ReachECO’s ability to successfully repair otherwise unsafe controllers in three continuous-control benchmarks: ACC, Docking2D and Car2D
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.