CoRe-WAM: Condition-Aligned Distillation and Recurrent Action Refinement for Efficient World-Action Models
Abstract
World-Action Models (WAMs) integrate world prediction with action generation, but iterative sampling increases inference latency during closed-loop control. Reducing this latency is therefore important for timely action updates in response to new observations. In this paper, we focus on the sequential formulation, which first predicts future videos and then generates actions conditioned on them, rather than jointly sampling both modalities. Distribution Matching Distillation (DMD) offers a way to accelerate both stages by training few-step generators with distribution-level supervision. However, a direct branch-wise DMD adaptation that retains ground-truth video conditioning for action matching introduces a mismatch with the student-generated futures used at inference. Moreover, video distribution matching alone provides no explicit action-level feedback for updating the predicted futures. To address these issues, we propose CoRe-WAM, a framework that combines Condition-Aligned Distribution Matching Distillation (CA-DMD) with Recurrent Action Refinement (RAR). CA-DMD shares the same student-generated future across the action sampler and both the teacher and fake action score models. A vector-Jacobian product (VJP) further propagates the action distribution-matching gradient through the action sampler to the video generator, incorporating downstream action feedback alongside video distribution matching. During inference, RAR constructs a proposal from the preceding chunk's terminal action and refines the noise-perturbed proposal under the current video prediction. Experiments on RoboTwin 2.0, LIBERO, and LIBERO-Plus demonstrate competitive manipulation performance with significantly reduced inference latency. On RoboTwin 2.0, our method reduces the mean inference latency per recurrent chunk from ms for the original LingBot-VA teacher to ms on an NVIDIA PRO 6000 GPU, yielding a speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.