SHARP-WAM: Sparse Hierarchical Attention with Probe Tokens for Bounded-Context World Action Modeling
Abstract
World Action Models (WAMs) present a compelling framework for robotic manipulation via video prediction. However, scaling WAMs to long-horizon tasks is hindered by the dense attention complexity of foundation backbones. Existing approaches either compromise non-Markovian dependencies via truncated sliding windows, or introduce external memory plug-ins that risk complicating the pre-trained backbone. Leveraging the insight that long-horizon action modeling relies on sparse historical interactions, we propose SHARP-WAM, a parse ierarchical ttention framework with Leanable robe Tokens for bounded-context World Action Models. This framework introduces a compact set of learnable probe tokens serving as contextualized descriptors and routers. Historical key-value states are retained within a native, fixed-capacity KV Set via online utility-driven eviction, thereby enabling horizon-invariant attention inference cost. Building upon this bounded backbone-native memory, we devise a hierarchical sparse attention scheme. Temporal Top- routing is first performed via probe tokens to select relevant candidates from the Fixed KV Set, followed by camera-aware spatial partitioning. This dual-level sparsity substantially curbs long-horizon computational overhead while preventing cross-view feature corruption and preserving fine-grained semantics. Extensive experiments demonstrate that SHARP-WAM outperforms existing methods on long-horizon tasks, achieving success rates of on RMBench, on LIBERO-Mem, and in real-world trials, while remaining competitive on general-purpose benchmarks. [Project page](https://anonymous.4open.science/w/SHARP-WAM_page/)
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.