acceptodds
Under review as a conference paper at ICLR 2027

SR: Fast and Faithful RL Rollouts via Self-Sparse Speculative Decoding

Abstract

Fast rollout generation for long-context reinforcement learning with verifiable rewards (RLVR) requires reducing KV-cache traffic without changing the response groups used for policy optimization. In our profiled RLVR workloads, attention accounts for 72–76% of autoregressive decoding kernel time. Sparse attention reduces these reads, but sampling directly from a sparse policy changes the response groups used for learning. For group-relative optimization, we show that even exact response-wise reweighting does not generally recover the expected dense-policy update because the advantage of each response depends on the rewards of other group members. We propose Self-Sparse Speculative Rollouts (\modelname). It uses sparse attention to propose tokens, then generates responses through full-attention verification and exact speculative sampling. Drafting and verification share the same policy parameter, eliminating the need to train or update a separate drafter. The resulting response groups follow the dense-policy distribution. Under \modelname, we characterize speedup conditions and the costs of drafting and verification. We found that dense verification still requires the complete historical KV cache. A KV-aware scheduler therefore limits KV-cache repeated offloading and restoration while maintaining rollout parallelism. % OR thrashing In a matched ablation with an aggressive sparse drafting budget, \modelname maintains stable training while direct sparse sampling with training-side correction becomes unstable. On long-context mathematical reasoning, \modelname achieves the end-to-end training throughput of autoregressive rollouts on Qwen3 model, with comparable task performance. Our code is available at https://anonymous.4open.science/r/S3R-7666

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.