AUGUstus: Autonomous Driving VLA Model with Unified Generation and Understanding
Abstract
Vision-language-action (VLA) models have shown strong potential for end-to-end autonomous driving by leveraging world knowledge and understanding to achieve great performance in long-tail scenarios. However, current VLAs suffer from two major limitations: 1) Degeneration into vision-action (VA) models. To achieve low latency, regression-based and flow-based methods directly generate actions, yet they do not sufficiently leverage the VLM's understanding, leading to unsafe and unexplainable actions. 2) Autoregressive high latency. Autoregressive methods sequentially generate chain-of-thought (CoT) and action tokens, fully leveraging the VLM's understanding, but high latency makes deployment difficult. To enable low-latency deployment while properly using CoT to assist action generation, we propose , an utonomous-driving VLA model with nified eneration and nderstanding. Specifically, we first unify diffusion-based action generation and autoregressive CoT understanding in a single model, where understanding is optimized only to serve generation. More concretely, joint CoT supervision preserves the VLM's spatial understanding in the shared representation and keeps it useful for safer and more accurate trajectories. Second, we propose AUGUstus-mode, which orders the action generation before the CoT so that the CoT supervises the reliable action during training while inference needs only the action output, enabling low latency. Finally, we introduce , a unified GRPO algorithm that jointly optimizes the action and CoT, further improving action performance. Extensive experiments demonstrate that achieves 94.6 PDMS on NAVSIMv1 and 90.7 EPDMS on NAVSIMv2, establishing state-of-the-art (SOTA) planning performance without scorer module. On nuScenes, our method achieves an average L2 error of 0.27 m and an average collision rate of 0.05, both the best among the compared methods. To validate the practicality of , we train it on our in-house dataset using 512 NVIDIA H800 GPUs for approximately two days, then evaluate it on an internal 1,000 km dataset and preliminarily deploy it on a real vehicle, achieving 12 Hz operation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.