acceptodds
Under review as a conference paper at ICLR 2027

REGAIN: Restoring Token Distributions in Low-Bit Language Models via Truncated-Support Distillation

Abstract

Low-bit weight quantization substantially mitigates the memory footprint and latency bottlenecks of large language models, yet aggressive compression regimes incur non-negligible capability degradation. While low-rank quantization error compensation effectively recovers coarse benchmark accuracy with minimal parameter overhead, it fails to restore the nuanced output distributions of the full-precision teacher. Specifically, even when Top-1 predictions appear intact, quantization systematically distorts the probability margins and relative rankings among competing tokens, precipitating severe compounding errors in sequential and structured generation tasks. To resolve this discrepancy, we propose Regain, a parameter-efficient recovery framework based on truncated-support distillation. Instead of optimizing over the entire vocabulary or naively mirroring teacher targets, Regain dynamically constructs a compact token support from the union of the Top- candidates from both the full-precision teacher and the quantized student. This design preserves teacher-favored modes while directly penalizing quantization-induced spurious competitors, using an auxiliary Ghost token to aggregate the residual probability mass for bounded computation. With the low-bit backbone frozen, Regain exclusively trains low-rank adapters under this compact distributional objective. Extensive evaluations across Qwen3 architectures from 8B to 32B, multiple post-training quantization schemes including GPTQ and QuIP#, and aggressive 2 to 4-bit regimes demonstrate that Regain substantially improves distributional fidelity and downstream performance across question answering, tool invocation, and structured reasoning, setting a new Pareto frontier for low-rank recovery of ultra-low-bit language models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.