acceptodds
Under review as a conference paper at ICLR 2027

Understanding Training Stability in LLM Reinforcement Learning: A Survey

Abstract

Reinforcement learning has become a key paradigm for enhancing capabilities of large language models, enabling remarkable advances in mathematical reasoning, code generation, and long-horizon tasks. However, compared with supervised fine-tuning, RL optimization is substantially more unstable, frequently suffering from loss spikes and policy collapse. Despite extensive surveys on RL algorithms and alignment, training stability in LLM RL has not been systematically reviewed. This paper provides a unified survey of stabilization techniques from an optimization pipeline perspective. We categorize existing methods into two stages: gradient generation, covering data sampling, importance sampling, and credit assignment; and gradient propagation, covering gradient regulation and numerical/system alignment. Based on this unified perspective, we systematically review representative methods, analyze their underlying stabilization mechanisms, summarize their advantages and limitations, and identify open challenges and promising directions toward stable and scalable RL for LLMs. We hope this survey serves as both a comprehensive reference for future research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.