acceptodds
Under review as a conference paper at ICLR 2027

WAM-Adapter: Efficiently Adapting Video-DiT World Representations to Robot Actions

Abstract

Pretrained Video Diffusion Transformers encode rich spatiotemporal priors from large-scale videos, making them promising backbones for robot control. However, efficiently adapting them into World Action Models remains challenging. Existing approaches reduce video-adaptation or action-prediction costs, but it remains unclear whether representations produced during future-video adaptation can directly support lightweight action prediction. To address this gap, we propose WAM-Adapter, a parameter- and training-efficient framework that freezes the Video-DiT backbone, adapts it with LoRA, and reuses multi-layer representations for action prediction. A lightweight readout organizes multi-layer K/V representations from current-and-future processing into Current, Early-Future, and Late-Future groups, applying group-specific cross-layer fusion for action-chunk prediction. Inference requires a single Video-DiT forward over the current observation and noisy future slots, without future-video generation or iterative action denoising. Without embodied pretraining, WAM-Adapter uses only 121M trainable parameters with a compact Wan-1.3B backbone, achieving 89.0% success across 50 RoboTwin 2.0 tasks, 98.5% on LIBERO, and 75.17% on LIBERO-Plus. To our knowledge, this is the smallest reported trainable-parameter budget for directly trained WAMs without embodied pretraining. Under the same setting, WAM-Adapter improves training throughput by 9.66 while using 98% fewer trainable parameters than Fast-WAM. These results show that Video-DiT representations formed during world-model adaptation can be efficiently reused for strong robot control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.