SEAM: Smooth Execution of Action-Chunked Motion for Vision-Language-Action Policies
Abstract
Vision-Language-Action (VLA) policies generate locally coherent action chunks, yet consecutive predictions can disagree in their overlapping future actions, introducing abrupt changes at execution boundaries. We propose SEAM (Smooth Execution of Action-chunked Motion), a training-free inference method for flow matching VLAs. SEAM uses the previous chunk's unexecuted tail as a temporally aligned consistency reference. Its core mechanism, Velocity-guided Loss Steering (VLS), alternates learned velocity-field steps with closed-form corrections toward a time-dependent mean reference. The corrected state feeds subsequent velocity evaluations, coupling cross-chunk consistency with observation-conditioned generation without policy-network backpropagation. In a 1,300-episode comparison on LIBERO-10 with π₀.₅, SEAM reduces boundary action second differences by 28% and chunk transition discontinuity by 27% at near-baseline denoising-loop cost. Matched-state controls show that the aligned reference improves the continuity–success trade-off over zero or shuffled references. With SmolVLA on LIBERO-Spatial, the same guidance setting reaches 72.0% success versus 65.6% for baseline and 60.0% for RTC, with respective realized end-effector jerk reductions of 58.8% and 40.0%, at approximately 1% additional full-policy execution cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.