TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on a Consumer-Grade GPU with <1 GB VRAM
Abstract
Vision-language-action (VLA) models commonly adopt an LLM-centric pathway, processing visual observations and language instructions through a large language model before predicting robot actions. Although effective, this design incurs substantial computation and memory overhead. In this work, we introduce , a compact VLA architecture built on a direct mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design directly constructs task-conditioned representations while avoiding the overhead of an LLM-centered execution pathway. On LIBERO, TurboVLA achieves 97.6% average success with only 0.2B parameters, 31.2ms inference latency, and 0.9GB inference VRAM on a consumer-grade RTX 4090. Notably, a 0.4B TurboVLA achieves 88.06% success on RoboTwin 2.0, even matching or outperforming substantially larger VLA policies. These results demonstrate that the simple design of TurboVLA can achieve high performance without requiring an LLM-centric execution pathway, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code will be available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.