SpatialLAM: Latent Actions Need Spatial Structure
Abstract
Latent action models (LAMs) learn compact representations of visual transitions, enabling policies and world models to learn from videos with little or no action annotation. However, most existing LAMs represent latent actions as spatially unstructured tokens, leaving their correspondence to the locations of visual changes implicit. This provides limited structural support for reusing local transition content across positions, especially in tasks that require predicting or controlling similar local interactions at different image locations. To address this limitation, we propose SpatialLAM, which organizes latent actions on a 2D grid aligned with the observation plane and uses locality-biased attention to ground each token in a consistent visual neighborhood. We evaluate SpatialLAM in both policy learning and world modeling under matched data, optimization, and latent capacity. SpatialLAM improves policy success from 76.8% to 83.0% on LIBERO-Long, from 64.6% to 70.7% on out-of-distribution LIBERO-Plus, and from 45.0% to 66.7% across three real-world manipulation tasks. It also reduces world-model FVD by 16.9% on Something-Something V2 and 10.9% on Bridge V2. These results suggest that the utility of latent actions depends not only on compactly encoding visual transitions, but also on explicitly organizing their spatial structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.