acceptodds
Under review as a conference paper at ICLR 2027

Hindsight Self-Distillation: Teaching Robots to Act from the Present

Abstract

Robot demonstrations reveal the consequences of actions, but a deployed policy must act before those consequences are observed. How can hindsight shape what a policy sees before it acts? We introduce Hindsight Self-Distillation (HSD), which uses training-only future observations to shape the visual summaries used for action generation. A privileged teacher and a current-observation student share video and action experts, with each action branch receiving visual information exclusively through its own summaries. Ground-truth action supervision trains the teacher, which transfers both action-flow predictions and summary representations to the student. An interface-aware analysis separates teacher dependence on unavailable future evidence, information lost through the student interface, and student approximation error. At deployment, the student computes its visual summaries once and reuses their keys and values throughout action denoising. On LIBERO and RoboTwin 2.0, HSD achieves average success rates of 98.50% and 92.31%, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.