acceptodds
Under review as a conference paper at ICLR 2027

Auditing Dynamic Sensitivity-Selective Quantization For Low-Precision LLM Reinforcement Learning

Abstract

Reinforcement learning has become a critical post-training phase for large language models, where the rollout stage dominates the computational and memory overhead. Low-precision rollout reduces the cost of reinforcement learning for large language models but introduces a rollout-training mismatch that can destabilize optimization and even leads to collapse. Selective high-precision protection mitigates this mismatch, but the modules that benefit most from such protection can change as the policy evolves. Consequently, a protection set determined through one-time calibration may become suboptimal during training. We formulate selective protection as a time-varying allocation problem under a fixed high-precision budget, periodically assessing the sensitivity of each module to low-precision quantization and reconfiguring the protection strategy based on their current benefit-to-cost ratios. This allocation mechanism complements Quantization-Aware Training (QAT): while QAT adapts weights to low-precision computation, our method adapts precision assignments to the evolving module sensitivities. Experimental results demonstrate that combining dynamic allocation with QAT effectively controls the rollout-training mismatch throughout the evaluation window, whereas using QAT alone or QAT with static protection exhibits earlier signs of instability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.