KV4-RL: Unlocking Faster Asynchronous Rollouts for RLVR with INT4 KV Cache Quantization
Abstract
Asynchronous rollout improves the efficiency of reinforcement learning with verifiable rewards by mitigating long-tail generation. However, sustaining higher concurrency compresses the same context-processing work into fewer decode steps, increasing the per-step cost of KV-cache processing. We characterize this KV concentration and show that KV-cache quantization provides substantially greater acceleration than weight-only or weight-activation quantization under the resulting workload. We further find that INT4 KV quantization can degrade learning despite importance correction. To address this, we propose KV4-RL, which combines calibration-free INT4 KV quantization with selective INT8 protection of attention sinks and recent context. This bounded protection reduces sampler mismatch while retaining efficient execution over the compressed cache. In Qwen3-8B-Base GRPO training, KV4-RL achieves rollout and full-iteration speedup over BF16 across training, with rollout speedup reaching as responses lengthen. Under BF16 evaluation, final-policy mean accuracy across six math benchmarks remains within 0.5 percentage points of BF16-rollout training for both Qwen3-4B-Base and Qwen3-8B-Base.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.