acceptodds
Under review as a conference paper at ICLR 2027

AdaRollout: Accelerating Synchronous RLVR with Adaptive Parallelism Switching

Abstract

Reinforcement learning with verifiable rewards (RLVR) relies on large-scale rollout generation, yet its performance is often bottlenecked by synchronous execution that is highly sensitive to long-tail decoding latency. Stochastic decoding produces sequences with widely varying lengths, where a few stragglers can stall batch completion and lead to poor accelerator utilization. This issue is exacerbated under static parallelism: as short sequences complete, the workload drastically shifts from high concurrency to long-tail dominated execution, rendering the initial parallel strategy inefficient. In this paper, we present AdaRollout, a system for state-preserving rollout re-parallelization that dynamically reconfigures the execution topology during rollout without restarting in-flight requests. AdaRollout profiles batch-completion throughput offline and derives a lightweight runtime policy that uses unfinished requests as a proxy for the remaining workload while accounting for reconfiguration cost. It further enables efficient topology reconfiguration through architecture-aware weight resharding, cross-topology KV-cache migration, and pre-instantiated runtime resources. By dynamically reallocating resources to long-tail sequences, AdaRollout achieves substantial end-to-end training speedup over optimal static parallelism, while maintaining consistent training dynamics (reward RMSE ) and downstream performance (accuracy RMSE ).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.