OneRobo: Temporally Structured Distribution Matching for One-Step Robot Policies
Abstract
Diffusion and flow-matching based robot policies model continuous action distributions but incur inference latency from iterative sampling. Few step distillation reduce this overhead, yet existing methods overlook the within-chunk temporal structure and cross-chunk execution consistency required for robot control. We introduce OneRobo, a one-step distillation framework for robot policies that preserves temporal structure. To accommodate the distinct dynamics of arm motion and continuous gripper control, we adopt channel-specific structured generation: arm generation and its distillation directions are constrained to a low-frequency discrete cosine transform (DCT) subspace to suppress unnecessary temporal variations, while teacher supervision on gripper changes across timesteps promotes stability. To improve cross-chunk consistency, we incorporate aligned historical-action initialization and learnable adapters into training. These mechanisms form a differentiable action-generation pipeline, with distribution matching and auxiliary supervision jointly applied to the final structured outputs. Across 50 RoboTwin simulation tasks with FastWAM, OneRobo reduces action generation from ten steps to one, with only an drop in average success rate relative to the multi-step teacher and an inference speedup. Additional experiments on and multiple real-world tasks further demonstrate its broad effectiveness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.