WISH: Margin-Free Latent Control Barrier Functions via World-Model Imagination
Abstract
Safety filters learned in world-model latent spaces commonly use binary failure labels to define unsafe states. However, their performance depends on the choice of a margin function that grades safety over nominally safe states, which is difficult to obtain in latent spaces. We propose WISH, a latent filter that measures the robustness of a latent state by the minimum cumulative disturbance effort required to reach a learned failure set against a fallback policy. Theoretically, for the exact modeled dynamics and value function, we prove that the resulting filter prevents failure when the initial state has sufficient disturbance resistance and the cumulative disturbance effort remains below a given budget. We also derive finite-horizon failure-probability bounds under stochastic imagination. In practice, WISH trains the latent safety value function entirely in the imagination of the world model by a minimum cost effort adversary against a fallback policy. Across four pixel-based continuous-control tasks, WISH achieves better safety-performance tradeoffs than existing latent filtering baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.