Attacking Action Discriminability in Latent World Models
Abstract
Latent world models support model-based control by learning action-conditioned state transitions in representation space. A key requirement of such models is that their predicted dynamics faithfully reflect the consequences of different actions, allowing downstream planners to distinguish alternative future trajectories. Many existing attacks are tied to specific action sequences or trajectories, whereas multiple alternative action sequences may reach the same goal, making persistence across plan variations and repeated replanning challenging. We instead target the action dependence of the latent dynamics itself, aiming to produce perturbations that are less tied to particular action sequences and remain persistent across multi-step rollout and replanning. Specifically, our attack operates at two complementary levels: action sensitivity suppression and multi-step discrepancy propagation. For the former, we construct action ensembles and directly reduce the dispersion of their predicted latent states, weakening the model's sensitivity to action variations. To further suppress the propagation of action-induced discrepancies over future rollout steps, we derive a finite-horizon sensitivity bound and reduce the dominant sensitivity of the latent transition dynamics. Additionally, we steer the attacked predictions away from their clean counterparts to induce erroneous rather than merely insensitive predictions. The proposed attack is evaluated on multiple JEPA-style latent world models across two control environments, demonstrating that suppressing action sensitivity in latent dynamics can effectively disrupt future-state prediction and downstream planning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.