acceptodds
Under review as a conference paper at ICLR 2027

HIRE: Hippocampal Replay for Test-Time Agentic Reinforcement Learning

Abstract

Gradient-free test-time reinforcement learning (RL) adapts large language model (LLM) agents through experience without updating model parameters, avoiding the weight-access requirements and parameter-optimization costs of training-time RL. In this paper, we identify a discovery–consolidation gap in test-time RL: agents can discover high-return behavior yet fail to sustain the achieved returns in later rollouts. In existing advantage-based methods, each update reweights the same frozen policy using re-estimated advantages, which can reverse action preferences that previously led to success. Inspired by complementary learning systems theory, we observe that advantage-based adaptation statistically integrates past experience but lacks a complementary hippocampus-like mechanism for recovering successful behavior. We therefore propose HIRE (HIppocampal REplay), which uses replay to recover successful behavior, allowing the adapted policy to build on previously achieved progress rather than rediscover it. We prove that HIRE satisfies a return consolidation criterion, yielding an expected-return lower bound that is non-decreasing as the best verified return increases for a fixed exploration rate. Experiments on ScienceWorld and WebChoreArena show that HIRE achieves the best overall performance among the evaluated methods. Our code is available at https://anonymous.4open.science/r/HIRE-D54E.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.