acceptodds
Under review as a conference paper at ICLR 2027

When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control

Abstract

Pretrained world models, learned simulators that encode an observation into a latent state and predict how it evolves under actions, are beginning to be reused as off-the-shelf dynamics backbones for control. We study a supply-chain backdoor in which the attacker controls only the released checkpoint, while the victim trains on clean data and never sees the trigger during training. To our knowledge, this is the first checkpoint-only world-model backdoor that induces targeted actions through a downstream controller trained or instantiated using clean data. The attack encodes no explicit trigger-to-action rule. Instead, the poisoned model routes trigger-bearing observations into a chosen latent region and reshapes the local dynamics there, so that the victim's own optimization, through Dreamer-style actor training or MPC/CEM planning, re-discovers the attacker's target action on its own. Across several control tasks and trigger families, the strongest evaluated settings control every action dimension and reach up to 100% Step-ASR, while harder and transfer settings show partial target recovery. Across the main settings, poisoned checkpoints retain roughly 74-100% of baseline clean utility and can evade clean-data diagnostics in most evaluated cases. Matched clean-model controls separate the backdoor from the effect of the trigger alone, and mid-episode tests show that the action changes only while the trigger is present. Moderate trigger-blind repair can leave triggered failure intact, while stronger adaptation trades off against clean control. These results identify the released world-model checkpoint as an underexamined attack surface for downstream control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.