acceptodds
Under review as a conference paper at ICLR 2027

MONAD: One Shared Diffusion Vision-Language Model for Continuous Control

Abstract

Vision-language-action policies begin with a vision-language backbone that emits discrete tokens and must end in continuous control. Every current answer pays a different price. A second module holds action-specific parameters, and its gradients degrade the pretrained backbone; the common remedy blocks them at the cost of fixing the separation. A symbol vocabulary commits the trajectory to bins whose errors carry no fixed physical size. One that keeps the backbone and swaps its objective writes a second target into the representation that language modeling trained. We observe that none of these prices is necessary. A diffusion vision-language model already reconstructs masked spans inside its own hidden space, and a complete action vector is a span like any other. We introduce a method that represents each control timestep as one continuous latent token and reconstructs temporally ordered action blocks with the same masked-token procedure the backbone applies to language. Two position-wise projections carry the coordinate mapping and hold 0.45% of model parameters. Action experts in π₀ and π₀.₅ instead allocate 16.5% and 19.2%. Layer probes place the cross-time response inside the shared backbone. The method adds no action expert, tokenizer, or codec, and replaces no pretrained objective. On standard benchmarks, it reaches 98.15% on regular LIBERO, 85.9% on the 10,030-task LIBERO-Plus aggregate, and 4.142 average completed chain length on CALVIN ABC–D.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.