HIGAP: Bridging the Overhead Gap in History-Based Speculative Rollout for MoE RL
Abstract
Rollout generation is a major bottleneck in reinforcement learning (RL) for large language models. History-based speculative rollout can reduce this cost by reusing trajectories from previous training steps as drafts, but its rollout-stage savings do not necessarily translate into end-to-end speedup, especially for mixture-of-experts (MoE) models. Expensive actor verification, redundant old-log-probability computation, and actor drift can substantially offset the benefits of history reuse. We present , a history-based speculative rollout framework that bridges this overhead gap through full-pipeline optimization. HiGAP adaptively bounds historical draft lengths to regulate history reuse and verification overhead, leverages verification probabilities to partially bypass old-log-probability recomputation, and mitigates history staleness under actor drift through post-update probability caching and advantage-guided history updates. Across three MoE and one dense models, HiGAP consistently reduces end-to-end step time and improves training throughput while maintaining competitive training effectiveness. On Qwen3.5-35B-A3B, HiGAP reduces step time by 11.6% and improves throughput by 12.7% over vanilla GRPO, while improving the average score by 12.3 points across nine mathematical reasoning benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.