LaDyP: Phase-Aware Action Chunking
Abstract
Action-chunking policies predict several future actions from the current observation, so a single chunk often spans a grasp, a release, or the switch to the next sub-task. On long-horizon instructions these transitions are where most rollouts fail. We introduce LaDyP, a phase-conditioning interface for Flow-based chunk policies. An offline teacher labels demonstrations with recurring interaction roles, within-role progress, and control-boundary evidence; a causal estimator predicts a soft phase belief, local progress, and current and upcoming boundary likelihoods from observation history. These signals condition the action generator through a global token, chunk-position embeddings, and a boundary-gated residual velocity. We attach LaDyP to a Flow policy trained from scratch and to the pretrained SmolVLA, and evaluate both on all four LIBERO suites with 1,500 paired trials per suite and method. On LIBERO-Long, LaDyP raises success from 49.2% to 57.5% for the Flow policy and from 71.8% to 78.7% for SmolVLA, while gains on the short suites are small. Delaying the signals by 32 steps or replacing all of them with signals from another episode removes most or all of the gain, and failures during engagement, release, and sub-task switches decrease by 31% and 30% for the two backbones while approach and transport failures stay near their Base counts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.