acceptodds
Under review as a conference paper at ICLR 2027

Does Imitation Learning Preserve Expert Performance Across Task Execution Speeds? A Benchmark Study

Abstract

Imitation learning seeks to reproduce expert behavior and is commonly evaluated by task success rate, the fraction of evaluation episodes completed successfully. Matching expert success under nominal conditions does not establish how the expert–learner performance difference changes with task conditions. We study how the expert–learner performance difference varies with execution timing. Timing changes are expressed through a scalar speedup factor; a factor of two halves the scheduled durations. We introduce ParcelStow, a benchmark with three contact-rich manipulation tasks—parcel insertion, upright placement, and keyed peg insertion—that varies these durations while fixing geometry, physical parameters, and success criteria. We evaluate Action Chunking with Transformers (ACT), Diffusion Policy, and DAgger using success rates, task-stage completion, and failure diagnostics. On parcel insertion, ACT matches the expert's 100% at nominal timing but falls to 53% at the maximum demonstrated speedup factor, compared with 84% for the expert. On upright placement and keyed peg insertion, ACT has higher nominal success than the expert, but this ordering reverses beyond the demonstrated speedup factors. These results show that nominal task success does not establish preservation of expert performance under changed task conditions, with execution timing providing a concrete counterexample.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.