acceptodds
Under review as a conference paper at ICLR 2027

Quantize Once, Adapt in Place: Zeroth-Order Optimization with Detectable Signals

Abstract

Quantization reduces the memory required to store and deploy large language models, while zeroth-order (ZO) optimization avoids storing activations for backpropagation. Combining these two advantages still leaves a key question: which parameters should be optimized? We therefore study a simple constraint: keep the packed weights and quantization metadata fixed, and fine-tune only the floating-point model tensors retained in the checkpoint. We collect these tensors into the Quantization-Preserved Subspace (QPS) and propose Quantization-Preserved Zeroth-order Optimization (QPZO), which applies two-point ZO updates only in QPS. This enables adaptation without adding trainable modules or dequantizing the low-bit backbone for optimization. We support QPZO with theoretical analysis and empirical evidence. Specifically, (1) we show that an updateable output head retained in QPS yields a nonzero projected gradient whenever its aggregate gradient is nonzero, establishing the presence of task-relevant signal; and (2) because ZO accesses this signal only through finite forward-loss differences, we introduce Signal-to-Error Ratio (SER) to compare directional task-signal energy with measurement-error energy. Under the stated error bound, we prove that SER above one is the sharp threshold for guaranteeing positive expected alignment, characterizing when finite forward queries robustly preserve a useful task direction. Our theoretical analysis and extensive experiments demonstrate that QPZO offers a simple approach that substantially reduces GPU memory usage while maintaining highly competitive performance and stability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.