acceptodds
Under review as a conference paper at ICLR 2027

Temporal Self-Imitation Learning

Abstract

Long-horizon policies trained with reinforcement learning can still achieve high return through inefficient interactions, while rare efficient behaviors discovered during training may be forgotten. We argue that temporal efficiency itself provides a source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors through efficiency-weighted self-imitation learning. Across 30 long-horizon tasks spanning robot manipulation and interactive navigation, TSIL consistently improves learning efficiency, behavioral efficiency, revisitation of fast successful behaviors, and robustness to unstable training conditions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.