Dynamism of a Trajectory: Encoding Coarse Intent into the Source of Flow-Matching Policies
Abstract
Robot policies are often trained on a relatively small corpus of demonstrations of a few tasks. In such low-data regimes, generative policies do not need to learn how their embodiments move because reproducing demonstrated trajectories is enough to fit the data. As a result, such policies tend to be competent only along those routes. Therefore, to make best use of limited demonstrations, we separate coarse and fine motion by first fixing an orthonormal basis for the coarse part of an action chunk where the coordinates of the basis represent displacement over time, which we call the command. Then, we modify the initially drawn Gaussian noise by replacing its component along the coarse motion basis with the command. The command is constructed to be unchanged along the training path and absent from the regression target, leaving the flow to learn the residual, fine motion. While a learned head predicts the command by default, an external planner or human operator can override it, and a scalar input allows inference-time modulation of command adherence. We evaluate our method on a quadrotor platform on gate navigation and out-of-distribution maneuvers. On the demonstrated tasks, our policy improves success in simulation from 45% to 73%, where the images differ from those seen in training. Our method also enables flexible out-of-distribution command steering for drone maneuvers through and around gates, such as orbits and figure-eights, in simulation and on hardware. Finally, we show a preliminary demonstration in which an agent pilots the policy by approving or overriding its proposed commands.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.