acceptodds
Under review as a conference paper at ICLR 2027

CONMAN: Steering LLM Agent Actions via Contrastive World Model Manipulation

Abstract

Large language model (LLM) agents can enhance action planning by leveraging a world model (WM) to predict the resulting next state for each candidate action, which is then scored by an evaluator to guide action selection. However, this emerging paradigm introduces a new attack surface when WM training is outsourced to potentially adversarial third parties due to limited computational resources or domain expertise. In this paper, we investigate whether an attacker can poison a WM so that a downstream agent is more likely to select an attacker-specified target action at a prescribed trigger state, without accessing or modifying the agent, its inputs, planner, or the state evaluator. We address this challenge through contrastive manipulation (CONMAN) of the WM's training data. Rather than manipulating the next-state labels for the target action alone, CONMAN contrastively reshapes the predicted next states of both target and competing actions: target-action promotion assigns the target action next-state labels that appear favorable to the evaluator by indicating progress toward task completion, while competing-action demotion assigns selected competing actions unfavorable next-state labels indicating obstruction. These labels are automatically generated and selected to polarize evaluator scores while remaining plausible next states. We evaluate CONMAN across ALFWorld, ScienceWorld, and WebShop with four domain-specific target actions. CONMAN reaches 43.6–70.0% end-to-end attack success rate, and 70.0–88.2% success rate among episodes in which the target action is proposed, outperforming a promotion-only baseline, which reaches at most 20.0% and 50.0%, respectively. Meanwhile, on separate utility evaluation sets containing tasks unrelated to the target action, CONMAN preserves native-task utility, without degrading the task success rate for any target.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.