acceptodds
Under review as a conference paper at ICLR 2027

ID2P: Future-Guided Policy Learning via IDM-to-Policy Distillation

Abstract

Future observations reveal the consequences of executed actions, motivating existing methods to leverage them for robot policy learning. However, prior work primarily uses future observations for representation-level supervision, while direct action-level future supervision remains less explored. We propose IDM-to-Policy Distillation (ID2P), a training framework that combines action-level and representation-level future supervision for flow-based policies. ID2P first trains an inverse dynamics model (IDM) conditioned on current and future observations using the same architecture as the policy. The trained IDM then provides action supervision by mixing its predicted action velocity with the expert action velocity from expert trajectories, while its action expert hidden states provide complementary representation supervision for the policy. Under an ideal IDM, we show that the mixed action velocity objective preserves the expected gradient of expert-only supervision while achieving no greater target variance. ID2P requires neither future observations nor explicit future state prediction at inference time, preserving the policy architecture and inference cost. Extensive simulation and real-world experiments demonstrate that ID2P consistently improves task success rates across flow-based policies of different architectures and scales, including SmolVLA (0.45B) and (3B).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.