VETO: A Sound Safety Envelope for Reinforcement Learning under Distribution Shift
Abstract
Safety envelopes filter the actions of a reinforcement learning policy through a predicted cost bound, and their appeal is that the bound is checked at every step. However, we find that the natural envelope fails in two ways under distribution shift. First, its fallback path reuses actions certified under an older calibration or at a different state, so the envelope is unsound and can emit an action its own bound rejects. Second, its calibration is fitted once, so the coverage the bound promises drifts with the deployment distribution. We introduce VETO, which combines a re-check of every emitted action at the current monitor state, a single temperature refit online from a scalar drift statistic, and a margin on the ensemble spread. Soundness then holds for any cost model and is machine-checked in the Rocq prover, and coverage transfer reduces to a Kolmogorov–Smirnov bound on one scalar, because coverage is a threshold event, which is why one temperature replaces per-step quantile tracking. In a pilot on two Safety-Gymnasium tasks with three constrained-optimization base policies, VETO lowers the per-step violation rate in 58 of 60 evaluation cells. The naive fallback emits inadmissible actions on up to 75.2% of steps, and a conformal filter with the same ensemble and no re-check has a violation rate at or above VETO’s in 54 of 60 cells. The price is an intervention rate that we report cell by cell. A sound envelope with online calibration is therefore what lets a per-step guarantee survive shift.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.