acceptodds
Under review as a conference paper at ICLR 2027

AgentFrontier: Stabilizing Learning Frontiers in Reinforcement Learning for Active Reasoning Agents

Abstract

Active reasoning requires large language model (LLM) agents to interact with external environments and strategically acquire missing information across multiple turns. Reinforcement learning (RL) with outcome rewards has become a standard approach for training such agents, particularly through group-relative algorithms that compare multiple rollouts from the same prompt to form relative learning signals. However, we observe that RL training for active reasoning often exhibits a sharp collapse after rapid initial improvement. We trace this instability to learning-frontier non-stationarity: under binary outcome rewards, a prompt contributes effective gradient signals only when its rollout group contains both successful and failed trajectories. As easier prompts become mastered, they silently exit the effective gradient computation, causing the learning frontier to shrink and concentrate on a narrow subset of hard prompts. Continued optimization then over-adapts to this shifting frontier while leaving mastered capabilities without gradient protection. To address this, we propose \ourmethod, a simple framework that stabilizes learning frontiers during RL training. \ourmethod combines Frontier Retention, which anchors the policy on mastered prompts through lightweight behavioral regularization, and Frontier Stabilization, which constructs difficulty-aware batches to smooth frontier composition across updates. Across  active reasoning tasks, \ourmethod consistently improves training stability and final performance, achieving gains of up to  with improvements on  out of  reported metrics, while reducing reward variance by over . These results highlight learning-frontier stability as a key principle for building reliable LLM agents capable of active reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.