acceptodds
Under review as a conference paper at ICLR 2027

TASK-SENSITIVE LOW-RANK QUANTIZATION ERROR RECONSTRUCTION WITH PROPAGATION-AWARE COORDINATION

Abstract

Low-rank quantization error reconstruction adds limited correction capacity to a fixed quantized model. Allocating that capacity requires a measure of task value, while deploying the corrections requires accounting for their responses under quantized states. We propose Propagation-Aware Task-Sensitive Low-Rank Quantization Error Reconstruction (PATR). Stage I combines input second moments with observed-token output gradients, reconstructs residuals in an active input subspace, and allocates heterogeneous ranks under a common factor-bit budget. Stage II fixes those directions and ranks, measures four single-group responses per Transformer Block under a correction-free quantized prefix, and fits their amplitudes in a suffix-sensitive metric. The Stage-I spectral gains constrain amplitude changes. Full-model Kullback–Leibler (KL) divergence selects a shared deployment step, whose amplitudes are absorbed into the existing factors. Across seven Llama and Qwen models, GPTQ and round-to-nearest (RTN) backends, and W4A8/W4A4/W3A4, complete PATR reduces WikiText2 perplexity (PPL) in all 42 matched configurations. A separate equal-budget rank-allocation study improves all 12 pairs, with a mean relative PPL reduction of 1.233%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.