TIER: Fine-Grained Trajectory Intervention for Stable Agentic Reinforcement Learning
Abstract
Large language models can substantially improve their reasoning capabilities through multi-turn reinforcement learning with external tools, yet prolonged training often suffers from instability and performance collapse. A common stabilization strategy filters out trajectories exhibiting recognizable failures, making a binary decision: each trajectory is either discarded or allowed to contribute fully to the policy update. Such coarse intervention not only requires increasingly elaborate filtering criteria to balance stability against data retention, but also leaves the optimization influence of retained trajectories uncontrolled. We find that rollouts without obvious behavioral degeneration can produce unusually large gradients and exert disproportionate influence on the aggregate gradient. To address this issue, we introduce Trajectory Intervention via Eligibility and Rescaling (TIER), a fine-grained trajectory intervention framework. TIER uses permissive screening based on generic response-level statistics to remove clear degeneration without attempting to enumerate every possible failure mode. It then clips excessive gradient contributions from eligible trajectories before batch aggregation while preserving their update directions. This design enables three levels of intervention: discarding degenerate trajectories, attenuating overly influential ones, and fully retaining the remainder. On AIME24, TIER improves over the corresponding TIR baselines by 10.2 and 14.0 points on Qwen2.5-7B and Qwen3-4B-Base, respectively. Experiments in retrieval-based interaction further demonstrate TIER's effectiveness beyond mathematical tool use. Overall, our results show that stable agentic reinforcement learning benefits from controlling not only whether each trajectory contributes to learning, but also how strongly it does so.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.