Dual-Loop Agents: Human-Level Atari via Prior-Guided Self-Improvement
Abstract
General self-improving agents must turn prior knowledge and limited online experience into effective action. However, existing approaches expose a tension: reward-driven reinforcement learning (RL) may fail to discover long-horizon strategies, whereas foundation-model priors often leave primitive actions unspecified. We introduce _dual-loop agents_, a self-improvement framework unifying priors and online experience through two coupled RL loops for learning high-level strategy and primitive control, linked by a shared goal–outcome interface. We instantiate this framework as PREinforce: a frozen foundation model constructs and revises an editable strategic program from online feedback, while a goal-conditioned neural policy learns fine-grained execution through temporal-difference contrastive learning. With 100,000 online interactions per game, PREinforce is, to our knowledge, the first agent to exceed the human reference in per-game mean score on all 26 Atari 100k games.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.