VSAD: Valuation-Sensitivity Attacks and Robust Defense for Latent World-Model Planning
Abstract
Model-based reinforcement learning (MBRL) achieves high sample efficiency by planning in the latent space of a learned world model. However, planners select actions by ranking candidate trajectories according to their predicted returns, making them vulnerable to local valuation errors.We show that these errors can be amplified by the planning process and identify valuation sensitivity as a characteristic signature of fragile high-value trajectories. Specifically, trajectories that receive spuriously high predicted returns exhibit substantial changes in valuation under small perturbations of the input observation, which correspondingly alter the initial latent state. They are also frequently associated with abnormal control patterns, such as action jitter and excessive impact, whereas successful trajectories exhibit more stable valuations. We further demonstrate that this sensitivity can be exploited at test time. In the white-box setting, bounded latent-state perturbations manipulate trajectory rankings and promote unreliable action sequences into the planner’s elite set. In the black-box setting, NES-based observation perturbations directly alter the target planner’s trajectory selection and induce similar control degradation. Based on this finding, we propose a training-free, model-agnostic defense that re-evaluates each candidate trajectory under multiple perturbed observations and ranks trajectories using a conservative score combining mean predicted return, return variance, and action smoothness. The defense suppresses trajectories whose high scores depend on unstable value estimates, without modifying the world model or requiring additional training data. Experiments show that it substantially reduces the effectiveness of both white-box and black-box attacks while preserving normal performance in clean environments. These results establish valuation sensitivity as a mechanistic signature of fragile trajectories, an exploitable attack surface, and a practical signal for robust planning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.