SimpleWAM: Training-Free Sink-Aware Compression for Efficient World Action Modeling
Abstract
World-action models (WAMs) jointly model future visual dynamics and robot actions, but long visual contexts and growing historical key-value (KV) caches can make inference expensive. We ask which parts of a pretrained WAM's visual context are actually necessary for action generation. Profiling action-to-visual attention in Fast-WAM and LingBot-VA reveals recurrent sink tokens with persistent action dependencies, alongside visual evidence whose relevance varies with the current action query. Based on this observation, we introduce SimpleWAM, a training-free sink-aware compression framework with two strategies: KeepSink protects profiled sinks while selecting the remaining action-visible tokens, and SinkAge extends sink preservation to historical KV through time-to-live scheduling and global budget allocation. On Fast-WAM, KeepSink remains competitive under a reduced token budget, whereas removing the profiled sinks drops RoboTwin Suite-10 success from 75% to 3%. SinkAge retains performance close to the Full KV while reducing the active visual-KV context by 55.1–69.4%. Experiments across WAM backbones further show that sink sensitivity varies with backbone and compression budget. In real-world experiments, KeepSink and SinkAge retain performance comparable to or better than Full KV under reduced visual-context budgets. Overall, SimpleWAM reduces active visual-KV cost without retraining while largely retaining task success, including under growing historical context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.