QALF: Deployment Error as a Regularizer for Quantization-Aware Low-Rank Factorization
Abstract
Low-rank factorization reduces a network's parameter count, while quantization reduces the bits used to store each parameter. Combining them usually means fitting the factors in floating point and rounding them afterwards. Factors that reconstruct a layer accurately can therefore perform poorly at the precision used for storage. We propose QALF, which accounts for quantization during fitting. Under an independent-noise model, we derive each factor's expected contribution to layer-output error in closed form. Adding these contributions to the input-weighted reconstruction loss gives a first-order approximation to the error after quantization. The quantization model determines the regularizer's form, and the target bit-width sets its weight, with no separately tuned trade-off coefficient. On ResNet-18/ImageNet, QALF outperforms dense GPTQ at comparable stored sizes up to the dense 3-bit budget. At 4-bit weights and 8-bit activations, it improves accuracy over the input-weighted reconstruction baseline by 1.8–2.8 percentage points with GPTQ and 4.1–4.7 with round-to-nearest, the rounding rule the objective models. It also exceeds CP-ADMM on its native grid, the closest quantization-aware factorization method, by 0.95–4.43 points at 4 bits. QALF remains more accurate than a tuned component-norm penalty combined with an optimized gauge, despite having higher kernel reconstruction error. The same derivation extends to matrix factorizations: on MLP down-projections in six language models from 410M to 8B parameters, QALF lowers 3-bit perplexity by 1.0–5.4 against SVD-LLM; its factors lose 71–87% less to rounding. For normalized CP factors, we establish conditions under which the regularizer prevents diverging components and the objective attains a minimum.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.