QCRL: Quantization-Consistent Reinforcement Learning for LLM Reasoning
Abstract
Obtaining a capable W4A4 reasoning model through reinforcement learning with verifiable rewards remains challenging because direct W4A4 optimization suffers from rounding-induced reward instability and rollout-optimization policy mismatch. We propose Quantization-Consistent Reinforcement Learning (QCRL), a dual-view framework that combines two strategies, Quantization-Consistent Rollout and Quantization-Weight Dampening, to improve the final W4A4 policy. Specifically, the rollout strategy collects W4A16 completions from round-to-nearest and stochastic-rounding views of shared latent weights and calibrates their rewards through cross-view comparisons and within-view support to guide W4A4 policy optimization. The weight-dampening strategy stabilizes weights near their quantization centers after parameter updates. On GSM8K with Qwen2.5-3B, QCRL achieves 78.32% pass@1 under W4A4 evaluation, surpassing BF16 reinforcement learning followed by W4A4 post-training quantization by 2.28 percentage points and direct W4A4 reinforcement learning by 4.25 percentage points. These results support reinforcement learning for stronger W4A4 reasoning models while retaining compatibility with W4A16 rollout.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.