acceptodds
Under review as a conference paper at ICLR 2027

QCRL: Quantization-Consistent Reinforcement Learning for LLM Reasoning

Abstract

Obtaining a capable W4A4 reasoning model through reinforcement learning with verifiable rewards remains challenging because direct W4A4 optimization suffers from rounding-induced reward instability and rollout-optimization policy mismatch. We propose Quantization-Consistent Reinforcement Learning (QCRL), a dual-view framework that combines two strategies, Quantization-Consistent Rollout and Quantization-Weight Dampening, to improve the final W4A4 policy. Specifically, the rollout strategy collects W4A16 completions from round-to-nearest and stochastic-rounding views of shared latent weights and calibrates their rewards through cross-view comparisons and within-view support to guide W4A4 policy optimization. The weight-dampening strategy stabilizes weights near their quantization centers after parameter updates. On GSM8K with Qwen2.5-3B, QCRL achieves 78.32% pass@1 under W4A4 evaluation, surpassing BF16 reinforcement learning followed by W4A4 post-training quantization by 2.28 percentage points and direct W4A4 reinforcement learning by 4.25 percentage points. These results support reinforcement learning for stronger W4A4 reasoning models while retaining compatibility with W4A16 rollout.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.