acceptodds
Under review as a conference paper at ICLR 2027

Learn from Failure, Act with Memory: Reflective and Prospective Learning for Reinforcement Learning

Abstract

Exploration–exploitation coordination is a core challenge in reinforcement learning. Existing methods rely on static weighting or annealing variants, limiting state-wise exploration–exploitation balance. This causes premature convergence or ineffective exploration, impairing learning efficiency and performance. Essentially, the coordination problem is not about static allocation but an experience-based decision problem requiring agents to remember, reflect on, and leverage past successes and failures. Inspired by human retrospective memory and prospective anticipation, we propose an integrated adaptive exploration–exploitation framework, . On the exploitation side, the Retrospective Module distills shared patterns from successes and identifies weaknesses from failures for targeted improvement. On the exploration side, the Prospective Module combines coarse global search with fine-grained state discrimination via sparsity-constrained latent representations. We theoretically justify proposed exploration mechanism. Furthermore, we introduce a policy-level ensemble to adaptively coordinate the prospective and retrospective modules, enabling a exploration–exploitation balance based on demands. We prove the coordinated strategy outperforms non-integrated strategies. can be plugged into existing algorithms. We evaluate across 56 environments against 23 baselines in total, where our method achieves better average performance than baselines. Notably, it achieves up to a 15× improvement on Montezuma’s Revenge. improves time efficiency over baselines by 41.8% on average.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.