acceptodds
Under review as a conference paper at ICLR 2027

Multi-Repacking in Asynchronous RL Post-Training

Abstract

Long-tailed trajectory lengths leave many rollout workers, or replicas, lightly occupied while a few unfinished trajectories continue. We propose multi-repacking, which jointly consolidates unfinished trajectories across multiple replicas. Under exponential service, linear cache growth, and service-preserving migration, we derive finite-system bounds and explicit hierarchy sums for the expected cumulative time that replicas remain occupied during the draining tail. Under the stated joint scaling assumptions and an explicit packing-slack condition, we prove that multi-repacking has a strictly smaller leading term than hierarchical pairwise repacking. We further establish a uniform positive gap between the corresponding hierarchy sums that dominates the lower-order errors in sufficiently large systems. The analysis allows replicas to start at different times and to enter the tail at statistically dependent times; it also permits packing decisions to adapt to observed completions and accounts for replicas that empty naturally without migration. In experiments with 1,600 replicas whose start times are independently staggered, each processing 12,800 trajectories in 3,200 parallel slots, all three tested group sizes \(m=3,4,5\) reduce mean tail occupied time on both evaluated workloads. Across 40 independent workload instances comprising 160 policy executions, the reductions are 4.74–7.59% under exponential service and 6.34–10.56% under resampled output lengths from an empirical SWE-agent coding workload. All six paired mean improvements have positive pointwise 95% confidence intervals. The trace-based runs use the same policy parameters, reserve physical cache headroom, and exhibit no overflow. These finite-system results demonstrate multi-repacking's occupied-time benefit under both the exponential model and the empirical coding-agent workload.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.