acceptodds
Under review as a conference paper at ICLR 2027

One Step Is Often Enough: Accelerating VLA Action Heads with Training-Free Marginal-Aligned Warm Starts

Abstract

Vision-language-action models with a flow-matching action head spend several denoising steps on every policy call. We ask how few the head needs at inference, without training, and find that one step is often enough but not always. The same single step keeps or improves success on some tasks and collapses on others. The split is not visible in open loop, where one step returns an estimate of the posterior mean with a small residual of similar size on every task. It is decided in closed loop, where that residual is persistent and integrates in the state until it exceeds the loop's tolerance. A one-step head therefore needs a starting point that carries the executed motion and lies on the marginal it was trained on. Marginal-Aligned Warm Start () provides both without training, by shifting the previously executed chunk, re-noising it to the flow's training marginal, and taking one step. Across models, simulated benchmarks and a real robot, lowers the residual as predicted, recovers most or all of the success that one step from noise loses, and outperforms training-free baselines that spend several evaluations per call. Because the warm start adds little to a one-step call, cutting the steps cuts the time. On the action head runs 9.1 faster at an unchanged suite average.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.