acceptodds
Under review as a conference paper at ICLR 2027

Accelerating World Action Models via Action-Aware Asynchronous Caching

Abstract

Multi-view world action models enable robots to predict future interactions from multiple camera views, while introduce substantial inference cost by multi-view processing. Although training-free caching has shown promise in accelerating diffusion and video world models, existing methods largely treat multi-view representations as a unified sequence and fail to exploit the uneven dynamics across different camera views, limiting their efficiency under aggressive acceleration. In this work, we propose A3Cache: Action-Aware Asynchronous Caching, a training-free acceleration framework for multi-view embodied world action models, which transforms multi-view caching into a asynchronous computation paradigm by adapting computation according to their update demands. Specifically, we introduce view asynchronous computing paradigm, which independently evaluates cache validity for different camera views and computes only the active views while predicting the stable ones by cache. We further introduce world-action contribution-aware token delaying paradigm within active views, which jointly considers token cache deviation and changes in visual-to-action contribution to prioritize urgent local computation while delaying less urgent tokens. Extensive experiments on multi-view robotic action demonstrate that A3Cache achieves a better efficiency-performance trade-off than existing methods, delivering substantially higher acceleration ratio with minimal performance degradation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.