ZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank Subspaces
Abstract
Zeroth-order (ZO) optimization enables fine-tuning large language models when backpropagation is unavailable or memory-prohibitive, but existing methods typically perturb the full parameter space or rely on structured low-dimensional perturbations, still leading to noisy gradient estimates and limited optimization flexibility. We propose ZO-Act, an activation-informed ZO fine-tuning method that uses input activations to define a fixed low-rank parameterization. For each linear layer, ZO-Act computes a small activation basis once at initialization and optimizes only lightweight coefficient matrices using forward-only loss evaluations. This explicit coefficient-space parameterization reduces the effective perturbation dimension, enables direct use of standard momentum-based optimizers such as Adam, and naturally supports quantized LLM fine-tuning by keeping low-bit weights frozen. Our analysis characterizes a dimension-alignment trade-off: reducing the perturbation dimension improves ZO gradient estimation, while effective subspace optimization requires preserving sufficient gradient energy. We further show empirically that activation-informed subspaces capture substantially more gradient energy than random and weight-based alternatives. Across Llama-3-8B, OPT-13B, Qwen2.5-7B, Qwen2.5-32B, and INT4 Llama-3-8B, ZO-Act consistently achieves competitive performance against strong ZO baselines across language understanding, question answering, commonsense reasoning, and mathematical reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.