SCEVLA: SPARSE COLLABORATIVE EXITS WITH SPECIALIZED TOKENS FOR DYNAMIC VISION LANGUAGE-ACTIONMODELS
Abstract
Vision-Language-Action (VLA) models have become a cornerstone of embodied intelligence, but their large-scale architectures introduce substantial inference latency, limiting real-time robotic deployment. Existing dynamic VLA approaches alleviate this issue by allocating computation adaptively. However, dense exit supervision and conflicting updates to shared parameters can impair intermediate predictions. In this work, we propose SceVLA: a dynamic-depth VLA framework that enables efficient inference via sparse collaborative exits with specialized tokens. SceVLA introduces sparsely distributed exits at different backbone depths, where each exit is equipped with dedicated action tokens to specialize in predictions at its corresponding representation level. To encourage effective cross-depth collaboration, we design a collaborative attention mechanism that allows exits to exploit complementary information from different layers, while blocking gradients through cross-exit feature transfers. Furthermore, we develop a lightweight router that dynamically selects an exit for each input through unidirectional attention over exit-specific tokens and backbone representations. On CALVIN, SceVLA achieves a policy-query speedup with AvgLen decreasing from 4.42 to 4.36. On RoboTwin 2.0, it achieves a speedup with 72.7% success, compared with 85.5% for the full-depth base policy, revealing a benchmark-dependent accuracy–latency trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.