acceptodds
Under review as a conference paper at ICLR 2027

PICA-WAM: Perturbation Invariance via Cross-Appearance Self-Alignment in World Action Models

Abstract

World Action Models (WAMs) jointly learn video prediction and action generation, yet remain fragile under visual perturbations that alter appearance without changing the task semantics or required actions. We identify a key source of this fragility in the different representation required by the two objectives: video prediction favors fidelity to visual appearance, whereas reliable action prediction requires invariance to task-irrelevant appearance changes. Existing approaches introduce external pretrained representation encoders. However, external feature alignment do not explicitly enforce consistency between the representations induced by nominal and perturbed observations. To address this gap, we propose PICA-WAM, a cross-appearance self-alignment objective that promotes perturbation invariance in WAMs. PICA-WAM aligns earlier-layer representations induced by perturbed observations with stable targets from later-layer representations induced by the corresponding nominal observations, explicitly enforcing cross-appearance consistency during training with no inference-time overhead. Experiments across diverse benchmarks demonstrate consistent robustness gains under visual perturbations, without compromising nominal performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.