Rethinking Quantized LLM Fine-Tuning: An Adaptive-Grouping Alternating Zeroth-Order Algorithm in Continuous Mapping Space
Abstract
Fine-tuning quantized Large Language Models (LLMs) is crucial for model adaptation on resource-constrained edge devices, where zeroth-order (ZO) optimization has emerged as a memory-friendly alternative to first-order methods. Existing ZO fine-tuning methods typically focus on either perturbing discrete quantized weights or coarse-grained (per-tensor/channel) continuous scales. In this paper, we rethink this fundamental choice and recognize that while fine-tuning should ideally target the dequantized weights, the discrete nature of weights restricts optimization to a finite search space, and tuning coarse-grained scales is limited to a constrained continuous subspace that lacks the expressive capacity of the full parameter space. To mitigate these issues, we propose to break down scales and zero-points to a fine-grained level and fine-tune them in a continuous mapping space. Motivated by this, we propose CoZO, an adaptive-grouping alternating zeroth-order algorithm that optimizes these fine-grained parameters while adaptively balancing search-space expressiveness and estimation stability. Crucially, we provide a theoretical convergence guarantee for CoZO. Extensive experiments demonstrate that CoZO achieves superior stability and effectiveness, consistently outperforming discrete-perturbation and scale-tuning baselines under stringent memory constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.