Making Computation Count: Speculative State Recycling for Flow-Based VLAs
Abstract
Speculative inference accelerates flow-based vision-language-action models (Flow-VLAs) by bypassing vision-language model (VLM) prefill and full flow integration on speculative rounds. However, lightweight drafting and approximate verification can compromise action quality, while additional computation to improve reliability can erode the resulting speedup. Moreover, rejection can erase these savings by triggering full-path inference after the draft and verification costs have already been incurred. These limitations point to a broader inefficiency: speculative states are often discarded once they are no longer directly usable, even though they may still contain useful predictive information. We propose Speculative State Recycling (SSR), which improves this quality–efficiency trade-off by recycling previously computed states across changes in conditioning. Before refresh, retained deep VLM features serve as semantic memory for a shallow predictor, augmenting the current observation with prior high-level representations. The same shallow computation also produces a draft-independent Context Risk estimate of how much refreshing the conditioning would change the full model's actions, guiding execution and refresh. When refresh is needed, saved proposals initialize short fresh-conditioned Action Expert integration, reducing regeneration from noise. On LIBERO, SSR with Triton nearly matches the success rate of full-generation with compiled inference while achieving a speedup per executed action and 74.1% fewer full AE calls per executed action. On LIBERO-Plus, SSR improves mean success by 5.77 points over FLASH at only 0.15 ms/action additional cost. Our project website is available at https://ssracc.github.io/speculative-state-recycling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.