acceptodds
Under review as a conference paper at ICLR 2027

Imagine a Little, Then Coordinate and Rethink: Modality-Adaptive Asynchronous Denoising for Unified Diffusion Vision-Language-Action Models

Abstract

Unified discrete-diffusion vision-language-action (VLA) models bridge world modeling and robot control by jointly generating future visual observations and actions within a single multimodal Transformer, providing explicit visual foresight for robotic manipulation. However, these models treat multimodal outputs as a homogeneous generation process, denoising them synchronously under a single shared schedule and overlooking the inherent heterogeneity across modalities: future-visual predictions are not only generation targets but also evolving conditions for action generation. Our analysis reveals distinct recovery dynamics across modalities and sample-dependent action responses to visual recovery. Motivated by these findings, we introduce MAAD-VLA, a modality-adaptive asynchronous denoising framework that preserves modality-specific progression while adaptively coordinating vision and action. During training, MAAD-VLA independently samples visual and action noise levels, exposing the model to diverse combinations of cross modal corruption so that it learns to denoise even when the two modalities are at different recovery stages. At inference, a response-guided two-stage denoising process first establishes partial visual foresight and then adaptively initiates action denoising alongside continued visual recovery under modality-specific schedules. We further introduce Action-centric Reconsideration, which enables self-correction by selectively remasking committed action tokens whose confidence drops sharply as the visual context evolves. Extensive experiments in both simulation and real-world environments demonstrate the effectiveness of MAAD-VLA, notably achieving an average task length of 4.82 on the CALVIN ABCDD benchmark.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.