Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It
Abstract
Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for improving these capabilities. However, RL alone often leads to instability or limited gains in tool-use tasks. We identify a recurring failure mode, termed structural collapse, in which performance drops sharply as valid tool calls become malformed control-token sequences. We analyze this failure at the token level and find that a small set of control tokens appears repeatedly across trajectories, accumulating reward-weighted training signals. They accumulate greater absolute advantage per token type than ordinary tokens, and their probabilities change much more sharply before collapse. Changing the invocation format recovers some performance, suggesting that part of the tool-use capability remains accessible. We further show that masking the direct policy-gradient (PG) terms at control-token positions can mitigate collapse. Building on these findings, we systematically compare supervisory signals, including off-policy supervision, hint-based guidance, and erroneous trajectory supervision, under synchronous and interleaved training schemes. We also introduce process reflection supervision, which uses teacher-generated error analyses and related examples to address the model's evolving failures. Across the evaluated models, interleaving supervised fine-tuning (SFT) with RL often improves stability, although the most effective supervision strategy varies by model. These improvements do not consistently transfer to format and content out-of-distribution (OOD) evaluation. Further analyses examine learning-rate effects and generalization across settings. Together, our findings identify a concrete failure mode in tool-use RL, provide token-level evidence about its development, and show how supervisory signals affect stability and generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.