How Deep Should a VLA Think When Thinking Costs Time? Budget-Constrained RL for Early Exit
Abstract
Vision-language-action (VLA) models are becoming capable generalist robot policies, but they are large, and their inference is slow on robotic hardware. Methods that make VLAs efficient are evaluated with the simulator paused while the model computes, as if the world waited. On a robot it does not: computation is consumed during control, and the value of an extra layer depends on how much control time it costs. We show in a controlled intervention that the same inference latency lowers task success by 3 to 40 points depending on how it reaches the robot, and we therefore ask how deep a VLA should think when thinking costs time. We cast adaptive depth as decision-making under a budget on the fraction of time spent inferring, and learn, with budget-constrained reinforcement learning, a lightweight policy that decides at each intermediate exit of a frozen VLA whether to act or to keep computing. Trained once per task suite in a simulator that keeps running while the model thinks, it runs \SPEEDFULL faster per inference than the full model on LIBERO-Plus, trading success for time as the budget sets, is on par with early-exit rules calibrated separately on every task at the same compute without any per-task calibration, and is ahead of an AVA-VLA-style exit gate on every task where their compute matches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.