acceptodds
Under review as a conference paper at ICLR 2027

When to Trust Your Teacher: Uncertainty-Aware On-Policy Distillation for Embodied Policies

Abstract

On-policy distillation mitigates compounding errors by aligning supervision with the student-induced state distribution, but its effectiveness is limited by a mismatch between this distribution and the teacher's reliable support. We study how to effectively leverage teacher policies for on-policy learning of flow-based robotic policies. We formulate OPD as policy optimization in a teacher-induced reward MDP and derive an uncertainty-penalized objective that, under an admissible uncertainty estimate, lower-bounds the return in the optimal-policy-induced MDP. Our analysis shows that optimizing this pessimistic objective yields a policy that is near-optimal among policies with bounded cumulative reward uncertainty, revealing a trade-off between the scope and tightness of the guarantee. We propose Uncertainty-Aware On-Policy Distillation (UA-OPD), a framework that accounts for potentially unreliable teacher supervision through uncertainty penalization. We instantiate UA-OPD for flow-based robotic policies with a self-consistency uncertainty estimator, combining soft penalization with uncertainty-based trajectory truncation. Experiments on LIBERO and RoboTwin 2.0 show that UA-OPD achieves a 1.5 improvement in sample efficiency and higher average peak success rates than OPD, while surpassing the teacher on most evaluated suites and tasks. Its benefits extend to out-of-distribution generalization and one-step policy distillation. These findings suggest that embodied OPD benefits from adapting the use of teacher supervision to its reliability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.