acceptodds
Under review as a conference paper at ICLR 2027

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become a powerful recipe for improving large language models on reasoning tasks, but its training cost is increasingly dominated by long-context rollout generation, making sparse attention as the rollout worker for training dense policy promising. However, in practice, mastering the tradeoff between training stability and efficiency is difficult. In order to study the optimal tradeoff, we first observe that the majority of the tokens are perfectly aligned with dense under highly aggressive sparsity. We then hypothesize that once the tail distribution is constrained by a threshold, we are able to train sparse rollout stably. We prove that our hypothesis holds by introducing a dynamic sparsity scheduling method for controlling the tail distribution constant through generation, and study how the threshold scales with model sizes. Surprisingly, we find that the threshold of 5-percentile mismatch at 0.86 generally works across model sizes and uses cost model analysis for finding the best scheduling for speedup. Empirically, our method generalizes to Qwen3-14B and enables stable training with sparse rollout, while achieving 2.2x, 2.4x, and 2.0x in rollout when training Qwen3-1.7B, Qwen3-4B, and Qwen3-8 B.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.