acceptodds
Under review as a conference paper at ICLR 2027

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

Abstract

Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation, and sampling-heavy evaluation. Unlike online serving, which typically targets request-level latency and throughput, synchronous rollouts must collect a complete batch before proceeding. A few long-tail generations can therefore dominate an entire rollout step, and uniform routing can slow these generations further by placing them in high-concurrency decoding batches. We present TailSieve, a partial-rollout-guided framework that jointly controls tail routing and replica allocation. Motivated by the observed persistence of prompt-level tail tendencies across policy updates, TailSieve uses partial-rollout completion status as a training-free signal for identifying candidate tail groups. Selected prompts are combined with newly sampled prompts, and all responses are generated from scratch under the current policy. A hierarchical controller uses response-work history and a serving model: a fast inner loop adjusts the number of isolated groups to balance completion times, while a slower outer loop adapts how many replicas serve these groups. The resulting low-concurrency execution also enables route-specialized speculative decoding with MTP or DFlash. Across the evaluated configurations, TailSieve achieves up to 1.67× speedup with routing alone and 2.59× with routing and speculation combined, both over uniform group routing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.