acceptodds
Under review as a conference paper at ICLR 2027

SALTQ: Saliency-Allocated Low-Bit Trainability for Quantized LLM Fine-Tuning

Abstract

Fine-tuning for deployment below 4 bits should directly produce a uniform low-bit weight-only checkpoint of integer codes and per-group quantization parameters, without an adapter branch or FP16 columns. Existing approaches fall short in different ways: adapter-based methods require post-merge re-quantization, mergeable variants freeze integer codes, weight-updating QAT computes the full weight gradient, and selective methods keep salient columns in FP16. To address this gap, we propose SALTQ, which allocates trainability by saliency within the deployed representation. A second-order analysis of the task loss shows that the columns where rounding costs the most loss are also those where a fine-tuning update moves the output the most, namely the few input columns with the largest activation energy, so that repairing the quantization error and adapting to the task can share one set of trainable columns. SALTQ makes these salient columns fully trainable at the deployment bit-width. In every other quantization group it trains only the zero-point, whose gradient pools within the group and is therefore almost free, whereas training the scales or the integer codes would require the full weight gradient. A channel permutation gathers the salient columns into complete groups; it is shared within contiguous segments of layers rather than across the whole network, which leaves one to three residual-stream reorderings at inference. Controlled comparisons show that the salient columns need to be trained but not in FP16: training them at the deployment bit-width retains nearly all of the gain of training them in FP16, with a gap of at most 0.8 points. At INT2 on Llama-2-7B and Qwen2.5-32B, SALTQ outperforms the compared mergeable and mixed-precision baselines by at least 0.7 points on the Commonsense-170k average and 3.8 points on GSM8K. At INT3, it reaches the level of the FP16 QLoRA reference on commonsense reasoning and GSM8K. It trains in about the time of QLoRA and in less than a third of the time of LR-QAT, which updates all weights and reaches a similar accuracy. The code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.