acceptodds
Under review as a conference paper at ICLR 2027

Per-Task Retraining Noise in Robot Imitation Learning

Abstract

Imitation learning trains robot policies from human demonstrations that a person collects one episode at a time by teleoperating the robot. Because gathering them is slow and expensive, curation methods try to make that effort go further by scoring each demonstration and retraining the policy on only the highest-scoring ones. Across the six curation papers we audit, reported success-rate improvements run from 5 to 32 percentage points in simulation. Yet retraining a policy on unchanged data also moves its success rate, and none of these improvements was measured against that retraining noise. We train 701 policies on controlled demonstration subsets, spanning diffusion policy and ACT from scratch and the vision-language- action models SmolVLA and π0.5 fine-tuned. Each is evaluated on PushT, two bi- manual ALOHA tasks and two ten-task LIBERO suites, every batch pre-registered. Retraining on identical data shifts one LIBERO task’s success rate by a standard deviation of 20.4 percentage points and the ten-task average by 4.23, comparable to the reported gains. This per-task noise recurs in Meta-World, where one re- train of a small state-based policy solves a drawer-opening task 200 times out of 200 and another solves it 0 times. No score we built for a demonstration predicts whether keeping it improves the policy, so the safeguard is to measure the noise itself. Part of it comes from limited evaluation episodes and is captured by evaluating one policy many times. The rest comes from training itself and appears only when training is repeated, so eight retrainings of one fixed demonstration set are enough to capture it. This measurement tells a lab whether a new policy is a real improvement or a lucky retrain. We propose that every curation gain be quoted beside the retraining noise of its own setting, and we release the run records and pre-registrations that make that comparison possible.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.

Per-Task Retraining Noise in Robot Imitation Learning | acceptodds