acceptodds
Under review as a conference paper at ICLR 2027

Lute: Folding Visual Hierarchies Around Action for World Action Models

Abstract

World–action models generate actions and predict the resulting future world states. Existing methods reduce future generation or reuse predictive features to make control more efficient. These strategies can still leave dense visual processing on the action path, while downsampling prediction targets removes spatial detail needed for world estimation. We introduce Lute (distinct strings, one shared resonant body), where folding visual hierarchies around action formation keeps action readout compact while preserving spatial detail for future prediction. Computation proceeds from visual analysis to action generation and then future prediction. Hierarchical processing progressively compresses visual features, with action tokens introduced only at the most compact stage. A paired analysis and synthesis transform retains complementary spatial information that bypasses action formation and supports subsequent prediction conditioned on action representations. All stages use the same block design, with roles defined by spatial scale and information access. Joint flow learning lets both objectives shape the shared hierarchy. When only actions are required, inference omits the subsequent prediction stages. Experiments show competitive control and improved tradeoffs between scene robustness and memory cost over efficient WAM baselines, while retaining structured future predictions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.