Learn from Failure, Act with Memory: Reflective and Prospective Learning for Reinforcement Learning
Abstract
Exploration–exploitation coordination is a core challenge in reinforcement learning. Existing methods rely on static weighting or annealing variants, limiting state-wise exploration–exploitation balance. This causes premature convergence or ineffective exploration, impairing learning efficiency and performance. Essentially, the coordination problem is not about static allocation but an experience-based decision problem requiring agents to remember, reflect on, and leverage past successes and failures. Inspired by human retrospective memory and prospective anticipation, we propose an integrated adaptive exploration–exploitation framework, . On the exploitation side, the Retrospective Module distills shared patterns from successes and identifies weaknesses from failures for targeted improvement. On the exploration side, the Prospective Module combines coarse global search with fine-grained state discrimination via sparsity-constrained latent representations. We theoretically justify proposed exploration mechanism. Furthermore, we introduce a policy-level ensemble to adaptively coordinate the prospective and retrospective modules, enabling a exploration–exploitation balance based on demands. We prove the coordinated strategy outperforms non-integrated strategies. can be plugged into existing algorithms. We evaluate across 56 environments against 23 baselines in total, where our method achieves better average performance than baselines. Notably, it achieves up to a 15× improvement on Montezuma’s Revenge. improves time efficiency over baselines by 41.8% on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.