Support-Aware Shielding Restores Safety for Predictive World-Model Control under Shift
Abstract
A predictive world model is supposed to let an operator ask “what if” before committing to a change, and that promise rests entirely on the model knowing when it is wrong. We study what becomes of the promise when the world moves. Working in a setting where every planning decision can be replayed against a full-information oracle, we find a pattern that repeats across every model we train: the planners that do best are the planners that are most dangerous. The models that extract the most benefit under familiar conditions are precisely the ones that breach the operator's safety constraints once a crowd forms, and the uncertainty intervals a shield would have to trust quietly stop covering the truth exactly when they are needed. This failure is invisible from forecasting accuracy alone: models that are nearly indistinguishable on held-out prediction error behave entirely differently once their forecasts are allowed to drive actions. We then show that the remedy is not a better forecaster. It is a planner that keeps asking whether it still recognises the world. A controller that re-plans as it goes, replays its own model against what it has just observed, checks whether its forecast has wandered outside the region its training data covered, and declines to act when either test fails, recovers nearly all of the lost safety while surrendering only a small and bounded share of the achievable benefit. We prove two statements that explain the exchange - one tying how reliably shifted situations are recognised to how often constraints can be breached, one showing that the cost of declining to act is paid in proportion to how often it is declined - and we machine-check both. The lesson we draw is that safety under distribution shift is more cheaply bought with an online admission of ignorance than with a more confident model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.