acceptodds
Under review as a conference paper at ICLR 2027

ReSPRIT: Reset-Aware Imagination from Self-Predictive Representations

Abstract

Model-free agents such as BBF achieve strong performance on Atari 100K, yet use their predictive models only to improve representations. Can imagination from a compact predictive model improve an already strong off-policy learner? We introduce Reset-aware Self-Predictive Representation Imagination Training (ReSPRIT), which adds reward and continuation prediction heads to BBF's existing latent prediction model to form a compact world model. Short imagined rollouts provide auxiliary actor and value learning alongside updates from replayed experience. Reset-aware scheduling limits the influence of unreliable imagined targets after parameter resets, while reduced imagination-loss weights balance learning from real and imagined experience. On Atari 100K, ReSPRIT improves the human-normalized interquartile mean (IQM) score from for the matched baseline without imagination to , the highest reported canonical IQM among the compared agents. It also achieves a higher game-level IQM than EfficientZero V2, while requiring no lookahead search at inference. In our Asterix comparison, ReSPRIT completes training and evaluation in roughly one-third the time of EfficientZero V2. Ablation studies support the benefits of imagination and the importance of reset-aware scheduling and loss weighting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.