acceptodds
Under review as a conference paper at ICLR 2027

Adversarial Attacks on World Action Models via Manipulating Imagination–Action Causal Consistency

Abstract

World-Action Models (WAMs) guide action generation by predicting future observations, requiring causal consistency between imagined futures and the consequences of generated actions. We reveal imagination–action causal consistency as a new attack surface in WAMs, where attackers can manipulate the causal relationship between imagination and action to achieve adversarial objectives. However, directly manipulating this consistency is challenging because its evaluation requires environment rollouts outside the model's computation graph, making the causal objective non-differentiable with respect to the perturbation. To address this challenge, we propose Causal Attack via Gradient Geometry (CAGE), a gradient-geometry-based attack framework that manipulates causal consistency through differentiable surrogates of imagination and action. Specifically, we design two tailored gradient strategies: orthogonal gradient decomposition isolates the optimization directions of the two objectives to disrupt causal consistency, while equiangular descent coordinates their optimization toward a causally consistent adversarial target to hijack imagination and action. We evaluate our attacks on LingBot-VA across LIBERO, RoboTwin, and real-world settings. Future Decoupling and World–Action Hijacking reduce the average task success rate from 96.0% to 15.2% and 28.6%, respectively. These results demonstrate that manipulating imagination–action causal consistency enables effective and flexible attacks on WAMs, revealing a new attack surface in prediction-driven robot control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.