BATON: Passing Denoising Trajectories Across Replans for Parallel Sampling of Robot Policies
Abstract
Diffusion and flow policies built on large action transformers generate each action chunk through many narrow, dependent model calls. We present BATON, a training-free sampler that evaluates the sampling steps of a frozen receding-horizon policy in parallel. BATON is based on two designs. First, it transports the previous denoising trajectory into the next replan by shifting the action horizon, completing the tail, and restoring the current noise boundary. Second, it starts two parallel Picard rounds over the full ten-step schedule from this trajectory and combines them with a depth-one Anderson extrapolation that adds no model calls. Ten dependent transformer calls thus become two wider ones. Without assuming contraction, we bound the error of the raw parallel rounds linearly in the initializer error, which is bounded by the change of the policy's plan between replans and the drift of its clean predictions, and explain the Anderson step as averaging an oscillation of the raw rounds. On 450 paired LIBERO-Plus instances at a fixed execution stride, BATON reduces Fast-WAM's median action-expert latency by 2.80× and backend latency by 1.73×, and solves 311 instances versus 304 for serial sampling, a difference that is not statistically significant. Paired ablations show that removing either transport or Anderson extrapolation lowers success. The speedups carry over to Fast-WAM on RoboTwin 2.0 (2.74× action expert, 1.68× backend), and with π0.5 on LIBERO-Plus the compiled full policy call becomes 1.30× faster, again without a significant change in success.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.