acceptodds
Under review as a conference paper at ICLR 2027

BRACE: Error-Guided Mixed-Precision LLM Quantization Without Backpropagation

Abstract

Post-training quantization (PTQ) reduces the deployment costs of large language models (LLMs). Uniform-precision quantization remains a widely used approach, but it often incurs substantial quality degradation at ultra-low bit-widths. In contrast, mixed precision quantization offers a better accuracy–budget trade-off by selectively assigning precision. However, at ultra-low precision, a quantization error cannot be judged by its local magnitude alone. Its significance also depends on the stage-specific response it should induce. This observation motivates a common design principle for bit allocation, error compensation, and quantizer calibration: quantization error should not merely be minimized locally, but should guide the response taken at each stage. Our approach builds on this observation and introduces three key components: (i) Curvature-Conditioned Residual Transport (CCRT) allocates bits using a joint curvature-conditioned cost that combines the error committed at a group with the one-step correction burden transported to its immediate successor; (ii) Block-Anchored Ridge Compensation (BARC) pairs reference and quantized paths within each block to derive a ridge-stabilized correction for propagated error before applying GPTQ to the corrected center under a metric with a ridge term; and (iii) Residual-Feedback Grid Refinement (RFGR) converts the elementwise errors of curvature-weighted first-pass grids into weights for a second scale and zero-point fit at the assigned bit-width, with adjacent-bit error reductions providing the feedback signal. Built on a GPTQ backbone, BRACE uses only a small calibration set and local layer-wise statistics, without backpropagation or iterative retraining. Extensive experiments on language modeling and zero-shot commonsense reasoning across the Llama2, Llama3, Qwen3, and OLMo3 model families show that BRACE achieves state-of-the-art performance among comparable backpropagation-free PTQ methods under 2- and 3-bit weight-only quantization settings. Code will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.