Freshness-Aware Prioritized Experience Replay for LLM Reinforcement Learning
Abstract
Reinforcement learning (RL) for large language models (LLMs) is expanding from mathematical reasoning to agentic tasks that require costly, multi-turn environment interactions. On-policy training repeatedly collects fresh trajectories and typically does not reuse them across rollout iterations. Experience replay can reuse these interactions, but replay priorities alone may not reflect their relevance as the policy evolves. We introduce FreshPER, a freshness-aware prioritized experience replay method for LLM training. FreshPER combines a base priority with exponential decay in trajectory age to balance priority and recency in the replay distribution. It augments fresh-rollout updates with historical replay while preserving the underlying RL objective and the current-to-behavior policy-ratio correction. Across Sokoban experiments with PPO, GRPO, and a ROLL REINFORCE++ variant, FreshPER improves over priority replay: under PPO, it reaches 43.0% validation success, compared with 33.6% for priority replay and 3.1% for on-policy training. Further results on NQ Search, ScienceWorld, and visual tasks support freshness-aware replay as a practical and effective way to learn more from costly interactions across diverse settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.