Beyond Reconstruction: Fisher-Regularized Post-Training Quantization for LLMs
Abstract
Large language models must be compressed to be deployed under realistic memory and latency budgets, and post-training quantization (PTQ) is dominant because it needs no retraining. Existing PTQ methods minimize a per-layer reconstruction error that approximates perplexity, but this proxy is not aligned with downstream task loss. Since a deployed model is typically dedicated to one task, the most direct way to preserve its task accuracy is to draw the calibration data from that task, yet this does not close the gap. We argue that the mismatch lies in the objective rather than in the data, and we introduce a simple yet effective modification that augments the PTQ objective with a Fisher-information regularizer preserving the weight directions to which the loss is most sensitive. We instantiate the regularizer in two forms, a Task Fisher computed from a labeled task training split and a Generic Fisher computed from the generic calibration data already used by PTQ, the latter removing any dependence on task labels and yielding a single quantized model that serves all tasks rather than a separate quantization per task. The regularizer requires no architectural changes and integrates into any per-layer PTQ solver as a diagonal loading of the per-layer curvature. Across five LLMs from 1B to 70B parameters and 2–4 bit widths, Task Fisher and Generic Fisher improve zero-shot accuracy and perplexity to a comparable degree, showing that neither task labels nor per-task quantization is necessary. Adding the label-free Generic Fisher to standard PTQ solvers raises multiple-choice accuracy by up to 6.3 points and generative-task accuracy by up to 9.4 points over the unregularized baselines, and brings the same improvement to a state-of-the-art solver.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.