acceptodds
Under review as a conference paper at ICLR 2027

VLA-Duo: Periodic Cross-Inference Scheduling for Mobile Vision-Language-Action Inference

Abstract

Mobile SoCs are an attractive platform for onboard vision-language-action (VLA) inference in robots with tight space and power constraints. However, their limited compute capacity makes it difficult to generate action chunks fast enough to support high-frequency control. We present VLA-Duo, a periodic GPU–NPU scheduling system that overlaps successive VLA inferences to increase action-chunk throughput at the cost of higher observation-to-action latency. For a fixed task configuration, successive VLA inferences use the same operator sequence and tensor shapes, allowing the same execution plan to be repeated. We generate the plan offline from GPU–NPU performance profiles, accounting for inter-accelerator interference. At runtime, the scheduler follows the plan to coordinate the two accelerators and overlap successive inferences. Short-run evaluation with π₀ and π₀.₅ on Snapdragon 8 Elite and 8 Elite Gen 5 shows that our system achieves 36.1–78.3% higher throughput than ExecuTorch/QNN, our strongest external NPU baseline. Output periods of 394–499 ms support action throughput equivalent to 10.0–12.7 Hz control when five actions are executed per chunk. In separate LIBERO simulations, π₀.₅ achieves 94.5% average success under a ten-control-step observation-to-action delay using prefix inpainting without additional training or gradient computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.