Towards Dense Capability Recovery of Ternary Quantized LLMs
Abstract
Ternary quantization offers extreme compression, but the resulting perturbation is substantially more severe than conventional low-bit quantization. The key challenge is not merely reducing one-shot projection error, but finding a ternary representation that remains efficiently recoverable through subsequent training. We demonstrate that under matched knowledge distillation (KD), quantizers with vastly different initial errors converge to remarkably similar functional endpoints. Motivated by this, we introduce a calibration-free, energy-preserving quantizer as a stable starting point designed specifically to facilitate subsequent training. However, same-size KD restores broad post-trained capability while leaving substantial and highly nonuniform performance gaps across tasks. To address this, we first exploit test-time sampling to expose substantial residual headroom on several reasoning and coding tasks. We then apply quantization-aware reinforcement learning (RL) directly to the ternary policy, converting part of this headroom into stronger single-shot performance. We evaluate our approach across several model sizes: Qwen3-1.7B, 4B, and 8B. The first math-RL stage substantially improves single-shot performance, while broader gains transfer across knowledge, reasoning, instruction following, mathematics, code, and the unoptimized thinking mode. At 8B, the final 2.03-GB model achieves 99.3% mean benchmark-wise retention relative to BF16 in non-thinking mode and 94.1% in thinking mode, while decoding at 77.9 tokens/s on an M5 Pro. Together, these results support a recovery-centric view of extreme quantization: preserve a recoverable ternary representation, restore the model distribution, and then optimize the deployed ternary policy directly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.