SOAR: Temporal Structure-Guided One-Step Action Generation for Robotic Manipulation
Abstract
Efficient robotic manipulation requires temporally coordinated actions with low inference latency. Multi-step diffusion and flow policies progressively refine action trajectories but require costly iterative inference, whereas one-step policies gain efficiency at the expense of temporal refinement. We introduce SOAR, a structure-guided one-step generative policy that improves action quality while preserving one-step efficiency. SOAR predicts an observation-conditioned discrete cosine transform (DCT) representation of future actions, capturing temporal variations across multiple frequency bands to provide explicit structural guidance. Unlike velocity-parameterized one-step flow policies that recover clean actions indirectly from a predicted velocity field, SOAR directly generates clean action trajectories in action space. This aligns the model output with the executable control target and provides a direct interface for structural calibration. The predicted structure further corrects the generated trajectory according to their spectral discrepancy, without requiring an additional evaluation of the generative backbone. During training, the corresponding MeanFlow velocity is derived from the calibrated clean action and used for consistency learning, while an endpoint objective aligns training with the pure-noise condition used for one-step inference. Across 53 Adroit and Meta-World tasks, SOAR achieves 83.9% overall success with only 8.7 ms inference latency. When applied to the VLA model π₀.₅, SOAR further achieves 97.4% on LIBERO and 60.0% on RoboTwin 2.0. Real-world experiments and extensive ablations further validate its effectiveness and efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.