Sparse-WAM: Decision-Aware Sparse Attention for Efficient World Action Models
Abstract
World Action Models (WAMs) jointly model future visual states and robot actions, but their large visual token sequences and iterative diffusion inference incur substantial computational costs. Existing sparse attention methods for video generation primarily preserve visual fidelity, which does not necessarily preserve the interactions required for action decisions. We propose Sparse-WAM, a decision-aware dynamic sparse attention framework for efficient WAM inference. Sparse-WAM calibrates attention budgets across heads and generation stages offline and adapts them through lightweight online sampling. It then uses action-to-video attention to identify action-relevant visual anchors and video-to-video dependencies to recover their supporting context, constructing sparse visual routes within the allocated budgets. To translate sparsity into practical acceleration, Sparse-WAM reuses stable budgets and calibration statistics across adjacent steps while updating sparse connections, and combines block-sparse execution with token reordering and cost-aware fallback. We evaluate Sparse-WAM on LIBERO using closed-loop task success rate and end-to-end inference latency. [Results to be added: achieved speedup and the corresponding success-rate difference from the dense baseline.]
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.