Fine-Tuning Low-Bit Models with Gradients in Quantized Code Space
Abstract
Fine-tuning low-bit models aims to adapt a quantized model while keeping the final deployed checkpoint in the same low-bit form. This setting is practically important as it reduces memory and inference cost for storage and deployment. Under this constraint, adaptation becomes an optimization problem over quantization codes and scales. Existing continuous low-bit training is efficient, but it can be distorted by straight through estimation error or by post-quantization gap; discrete search is deployment-faithful, but can be computationally inefficient under a finite training budget. We propose code surrogate gradient to find the steepest direction of a local continuous surrogate in code space to accelerate optimization, and perform guided search to preserve deployment faithfulness. Experiments across arithmetic reasoning, instruction following, and structured language understanding show that GradCodes consistently improves fine-tuning low-bit models across different quantization datatypes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.