HETEROGENEOUS ACTION CACHING: FREEING THE ACTION SOLVER OF WORLD ACTION MODELS
Abstract
World Action Models (WAMs) are markedly more robust under distribution shift than present-only policies, yet they are costly to run. The common practice of storing the future representation once lets a single video pass serve every action step, but it leaves the dominant cost untouched: once the future context exists, the action solver absorbs 84% of the runtime (165.27 of 195.29 ms on FASTER-WAM), against 25.60 ms for the visual branch and 6.04 ms for VAE encoding. The solver has been neglected as a caching target at the granularity that matters: latent-space and block-level caches act on spatially resolved features, while action-policy caches reuse whole chunks or demand an offline schedule or a learned gate. We present HAC (Heterogeneous Action Caching), which carries curvature-style heterogeneous token caching into action space. It treats each of the 32 chunk positions as a token, derives a scale-free curvature of its denoising path from the three most recent full evaluations, and forecasts each skipped step with a token-specific rule—copy for stable tokens, first-order extrapolation for linear tokens, and a smoothstep-damped blend for chaotic tokens—while a chaotic-only drift budget decides when a genuine evaluation is required. HAC trains nothing, searches nothing offline, and adds no learned parameters. It removes six of ten solver evaluations, speeding up the pipeline by 1.63× (1.86×on the solver) over FASTER-WAM at η=12: +0.09 points on RoboTwin 2.0 and +0.35 on LIBERO (seed 42), and −0.22 points on the matched LIBERO-Plus OOD set. The drift budget η acts as a single speed–quality knob, lowering latency monotonically from 8 to 12 while leaving in-distribution success within noise.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.