acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Demonstration Unlearning in Imitation Learning for Robotics

Abstract

Robot policies learn from demonstrations, some of which may later need to be withdrawn, for example when consent is revoked or an episode proves faulty. Demonstration unlearning edits the trained policy to remove their influence; ideally, the result matches a policy trained without them. Unlearning is often judged by isolated proxies such as restored task success, higher loss on withdrawn data, or membership-inference scores. When these proxies disagree with retraining, an edit can be credited with removing demonstrations whose influence remains detectable, or rewarded for overshooting what retraining would produce. We propose a retrain-calibrated audit that makes removal claims testable and shows whether an edited policy departs from retraining in behavior, data fit, or both. It compares the edited policy with independently retrained policies through actions executed at the same observations and prediction error (loss) on withdrawn episodes, including whether that loss separates them from never-used episodes; a symmetric rank test calibrates these differences against retraining variation. In physical cup insertion, an ACT edit raises success from 5/20 to 18/20 trials, yet its loss still perfectly separates audited withdrawn from never-used episodes and its actions remain far from retraining; a second policy class shows the same success-versus-evidence disagreement. In simulation, action and episode measurements detect complementary discrepancies; tests on fresh retain-only models and controlled data reintroduction reveal false rejections and sensitivity to checkpoint selection and queried states. Useful robot repair is thus not evidence of demonstration removal: unlearning claims should be checked against retraining in behavior and data fit.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.