acceptodds
Under review as a conference paper at ICLR 2027

PATCH: Propagation-Aware Two-Stage Quantization Error Compensation via Fine-Grained Correction for High-Fidelity 2-Bit LLMs

Abstract

Extreme low-bit quantization reduces the memory and data-movement costs of large language models (LLMs), but 2-bit precision introduces substantial errors that accumulate across layers. Existing compensation methods are constrained by both correction expressiveness and optimization scope: limited-rank corrections struggle to capture the error structure, channel-wise scaling overlooks intra-channel heterogeneity, and locally optimal compensation strengths need not minimize end-to-end loss. We introduce PATCH, a propagation-aware two-stage framework that decouples fine-grained correction from layer-wise strength regulation. The first stage fits group-wise compensation using activations from the quantized network, capturing heterogeneous intra-channel errors without an explicit rank constraint. The second stage fixes the learned corrections and jointly optimizes layer-wise regulation coefficients against an end-to-end objective, accounting for error propagation while preserving the correction directions. The resulting coefficients are fused into existing quantization parameters, avoiding a separate compensation branch. Across Gemma-2, Llama-3, and Qwen3 models of varying scales under 2-bit GPTQ, PATCH reduces WikiText-2 perplexity by 24.7% on average and improves average zero-shot accuracy by 2.03 percentage points relative to the best competing compensation method for each model and metric. These results demonstrate that jointly addressing correction expressiveness and network-wide compensation strength enables effective accuracy recovery in extreme low-bit LLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.