acceptodds
Under review as a conference paper at ICLR 2027

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving

Abstract

Autonomous driving has progressed from modular pipelines toward end-to-end unification, and Vision-Language-Action (VLA) models are a natural extension of this journey beyond Vision-to-Action (VA). In practice, driving VLAs have often trailed VA on planning quality, suggesting that the difficulty lies in the interface through which semantic reasoning, temporal context, and continuous control are combined. We argue that this gap reflects how VLA has been built—as isolated subtask improvements that fail to compose into coherent driving capabilities—rather than what VLA is. We present MindVLA-U1, the first unified streaming VLA architecture for autonomous driving. A unified VLM backbone produces autoregressive (AR) language tokens (optional) and flow-matching (diffusion-style) continuous action trajectories in a single forward pass over one shared representation, preserving the natural output form of each modality. A full streaming design processes the driving video framewise rather than as fixed video-action chunks under costly temporal VLM modeling. Planned trajectories evolve smoothly across frames while a learned streaming memory channel carries temporal context and updates. The unified architecture enables fast/slow systems on dense/sparse Mixture-of-Transformers (MoT) backbones via flexible self-attention context management, and exposes a measurable language-control path for action: a language-predicted driving intent steers the action diffusion via classifier-free guidance (CFG), turning language-side intent into control signals for continuous action planning. On the long-tail WOD-E2E benchmark, MindVLA-U1 surpasses experienced human drivers for the first time (8.20 RFS vs. 8.13 GT RFS) with only 2 diffusion steps, achieves state-of-the-art planning ADEs over prior VA/VLA methods by large margins, and matches VA-class latency (∼16 FPS vs. RAP's ∼18 FPS at matched ∼1B scale) while preserving natural language interfaces for human–vehicle interaction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.