acceptodds
Under review as a conference paper at ICLR 2027

ElasticWAM: Adaptive Spatiotemporal Imagination for Efficient Robot Control

Abstract

World Action Models (WAMs) benefit from imagining the future, but must they refine every region and complete future-video denoising to generate effective actions? We argue that future refinement should be guided by its contribution to action generation rather than visual completeness. We introduce ElasticWAM, an action-guided framework for adaptive spatiotemporal imagination. Built on a frozen pretrained WAM, ElasticWAM learns a unified refinement-demand field from action-prediction supervision to determine which future-video tokens to refresh, how far to advance denoising, and when to stop world refinement. It selectively updates future-video representations while retaining cached context from unrefreshed regions through the pretrained action-conditioning interface. Experiments across four manipulation benchmarks demonstrate the effectiveness of this approach. On RoboCasa-GR1, ElasticWAM improves success by 2.4 percentage points while reducing inference latency by 32.2% relative to Joint-WAM. Real-world dual-arm experiments further demonstrate long-horizon task execution under nominal conditions, lighting perturbations, and object replacement. These results support maintaining useful future context for control through selective refinement and cache reuse, without refreshing every region at every denoising step.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.