RANA: Accelerating Multi-Turn Agentic Reinforcement Learning with Speculative Decoding
Abstract
The evolution of Large Language Models (LLMs) toward reasoning-intensive agents relies heavily on Reinforcement Learning (RL) for alignment, e.g., RL human feedback (RLHF). However, the RLHF rollout phase suffers from severe computational inefficiencies, particularly in multi-turn interactions where sequential dependencies exacerbate the long-tail effect and lead to GPU underutilization. While recent work has explored speculative decoding (SD) to accelerate RLHF rollouts, they overlook the structural properties of agentic multi-turn interactions, i.e., contextual continuity, entropy decay, and hardware sensitivity, leaving substantial optimization potential unexploited. In this paper, we propose RANA (Rollout Adaptive N-Gram Accelerator), a system that accelerates multi-turn agentic RLHF rollouts by systematically exploiting these structural properties. RANA consists of three innovative components: (1) a Tree-based N-gram Speculative Decoding mechanism that leverages historical context for model-free drafting; (2) a Load-Adaptive Controller that dynamically toggles speculation based on Roofline analysis to avoid compute-bound contention; and (3) an Entropy-aware Scheduler that leverages entropy decay to align low-entropy rollouts with the execution tail, where speculative decoding yields the greatest benefit. Experiments show that RANA achieves up to 3.74x speedup, significantly outperforming prior approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.