Tail-W4: Predictive Precision Scheduling for the Long Tail of RL Rollouts
Abstract
Reinforcement learning (RL) post-training can be dominated by rollout, where synchronous execution waits for a few long responses to finish. Weight-only 4-bit quantization (W4) is promising for addressing those small-batch tails because it reduces the cost of loading model weights. However, naïvely applying W4 in RL post-training does not guarantee faster RL steps: quantized sampling can generate more tokens, increasing both decoding and downstream RL work. To address this challenge, we propose TAIL-W4, a rollout precision scheduler that chooses when to switch from BF16 to W4 by jointly predicting response lengths and execution costs. Reserving W4 for the tail can mitigate token inflation, but switching too early risks extra work and switching too late misses acceleration. Fixed schedules require repeated trials to tune and cannot fully adapt to changing training workloads. TAIL-W4 calibrates estimates of token inflation, updates them from training rollouts, and periodically predicts the benefit of switching to W4. On Qwen3.5-4B, TAIL-W4 achieves 1.38× rollout and 1.26× measured RL-step speedup over BF16, versus 1.12× and 1.03× for uniform W4 with the same quantized weights.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.