VLA-Duo: Periodic Cross-Inference Scheduling for Mobile Vision-Language-Action Inference
Abstract
Mobile SoCs are an attractive platform for onboard vision-language-action (VLA) inference in robots with tight space and power constraints. However, their limited compute capacity makes it difficult to generate action chunks fast enough to support high-frequency control. We present VLA-Duo, a periodic GPU–NPU scheduling system that overlaps successive VLA inferences to increase action-chunk throughput at the cost of higher observation-to-action latency. For a fixed task configuration, successive VLA inferences use the same operator sequence and tensor shapes, allowing the same execution plan to be repeated. We generate the plan offline from GPU–NPU performance profiles, accounting for inter-accelerator interference. At runtime, the scheduler follows the plan to coordinate the two accelerators and overlap successive inferences. Short-run evaluation with π₀ and π₀.₅ on Snapdragon 8 Elite and 8 Elite Gen 5 shows that our system achieves 36.1–78.3% higher throughput than ExecuTorch/QNN, our strongest external NPU baseline. Output periods of 394–499 ms support action throughput equivalent to 10.0–12.7 Hz control when five actions are executed per chunk. In separate LIBERO simulations, π₀.₅ achieves 94.5% average success under a ten-control-step observation-to-action delay using prefix inpainting without additional training or gradient computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.