Rising Tide: Overcoming Learnt Exploration Avoidance
Abstract
We introduce Rising Tide, a new algorithmic approach for learning performant policies that, unlike reinforcement learning methods, does not explicitly maximise expected return. Given a base policy, Rising Tide learns a sequence of policies indexed by return threshold, each approximating the base policy restricted to trajectories above that threshold, and trains higher-threshold policies by bootstrapping from lower-threshold ones. This design addresses a failure mode we identify as Learnt Exploration Avoidance (LEA): the tendency of RL agents to actively suppress exploratory behaviour because it initially appears suboptimal. We show that, in LEA-inducing environments, LEA can prevent standard RL methods from reaching optimal behaviour even with effectively unlimited training time. Rising Tide is architecture-agnostic; we demonstrate it with Transformer, LSTM, and fully connected architectures. To evaluate LEA in standard RL, we introduce Tree Climb, a new benchmark designed to isolate LEA. On three established meta-RL domains, Rising Tide delivers consistently strong performance, and on Tree Climb it is the only method to solve either of the two hardest variants.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.