AgentFrontier: Stabilizing Learning Frontiers in Reinforcement Learning for Active Reasoning Agents
Abstract
Active reasoning requires large language model (LLM) agents to interact with external environments and strategically acquire missing information across multiple turns. Reinforcement learning (RL) with outcome rewards has become a standard approach for training such agents, particularly through group-relative algorithms that compare multiple rollouts from the same prompt to form relative learning signals. However, we observe that RL training for active reasoning often exhibits a sharp collapse after rapid initial improvement. We trace this instability to learning-frontier non-stationarity: under binary outcome rewards, a prompt contributes effective gradient signals only when its rollout group contains both successful and failed trajectories. As easier prompts become mastered, they silently exit the effective gradient computation, causing the learning frontier to shrink and concentrate on a narrow subset of hard prompts. Continued optimization then over-adapts to this shifting frontier while leaving mastered capabilities without gradient protection. To address this, we propose \ourmethod, a simple framework that stabilizes learning frontiers during RL training. \ourmethod combines Frontier Retention, which anchors the policy on mastered prompts through lightweight behavioral regularization, and Frontier Stabilization, which constructs difficulty-aware batches to smooth frontier composition across updates. Across active reasoning tasks, \ourmethod consistently improves training stability and final performance, achieving gains of up to with improvements on out of reported metrics, while reducing reward variance by over . These results highlight learning-frontier stability as a key principle for building reliable LLM agents capable of active reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.