acceptodds
Under review as a conference paper at ICLR 2027

ACU-RL: Adaptive Correction and Update Control for Low-Precision LLM Reinforcement Learning

Abstract

Low-precision rollout can reduce the dominant generation cost of reinforcement learning for large language models, but it also creates a policy mismatch between the quantized behavior policy and the BF16 learner. Existing truncated importance sampling (TIS) mitigates this mismatch by reweighting sampled tokens, yet it can strongly attenuate tokens favored by the quantized rollout policy. We find that relaxing this lower-side attenuation improves early learning but can amplify late-stage policy drift, revealing that sample correction and actor stabilization must be controlled jointly. To address this trade-off, we propose Adaptive Correction and Update Control for RL (ACU-RL), which couples Adaptive Two-Sided Truncated Importance Sampling (AT-TIS) with risk-gated actor-update control. AT-TIS adapts the correction strength while retaining more of the gradient contributions from tokens favored by the quantized rollout policy. Complementing this sample-side correction, risk-gated actor-update control increases drift regularization when persistent excess mismatch coincides with degradation signals and reduces the policy-gradient multiplier when the mismatch becomes more severe. ACU-RL requires no additional sampling or model forward passes. We evaluate ACU-RL on Qwen2.5-Omni-3B and Qwen3-8B-Instruct across six multimodal and four mathematical reasoning benchmarks, respectively. It achieves average accuracies of 45.26% and 60.82%, compared with 44.72% and 58.47% for BF16 RL. NVFP4 rollout further provides a end-to-end training speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.