acceptodds
Under review as a conference paper at ICLR 2027

MuxWAM: A Compact World–Action Model with a Neural Multiplexer

Abstract

World–Action Models (WAMs) extend action-centric Vision–Language–Action (VLA) policies with explicit foresight by jointly predicting future world states and robot actions. However, dense future-image prediction in WAMs incurs inference overhead, while recursive decoding and re-encoding can compound prediction errors. To address these challenges, we introduce MuxWAM, a compact WAM whose unified multiplexer-inspired world constructor (MUX) maps dense visual patches to a fixed set of task-conditioned world tokens per view. MUX uses task-indexed routing over shared experts to refine visual summaries, with normalized global fusion completing the construction. The model jointly predicts future world tokens and actions through flow matching, without future-image decoding at inference. We evaluate MuxWAM on LIBERO, LIBERO-Plus, Meta-World, and physical ALOHA manipulation tasks. With only two world tokens per view, MuxWAM achieves 98.8% mean success on LIBERO, 83.76% averaged equally over Meta-World's four difficulty groups, and 78.3% across three physical tasks under normal conditions. These results demonstrate the effectiveness of task-conditioned compact world representations for robot control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.