RECAP: Refine Quantized LLM Checkpoints, Don't Replace Them
Abstract
When a deployed quantized large language model requires higher accuracy, transmitting a complete replacement checkpoint repeats information already available at the client. We formulate successive model refinement as encoding an incremental correction with the cached checkpoint as shared side information. The central challenge is to recover model quality with an incremental correction substantially smaller than a replacement checkpoint. We introduce RECAP, which combines two correction representations with different communication costs. An activation-subspace code transmits coefficients in a basis reconstructed at both endpoints from the cached model and public calibration data, with water-filling allocating bits by activation-weighted coefficient energy; a complement code transmits low-rank factors to correct the remaining residual. Our analysis establishes the shared basis's minimax optimality, derives optimal coefficients and approximate rate allocation, and quantifies the complement's benefit. We ask whether a correction can deliver at least the quality of the next-higher-bit replacement checkpoint while transmitting fewer bytes. From four 1.58-bit checkpoints of 7B models, RECAP reaches the perplexity of same-family 2-bit replacements with 2–19% of their bytes; at a fixed budget, it also achieves lower perplexity than conventional reconstruction baselines at both 7B and 13B scale. These findings show that even a checkpoint with poor standalone performance can remain valuable side information for communication-efficient model refinement. Code is available https://anonymous.4open.science/r/RECAP-7F3A/README.mdhere.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.