Stabilizing Zeroth-Order Fine-Tuning of Mixture-of-Experts via One-Sided Estimation and Expert Replay
Abstract
Zeroth-order (ZO) optimization enables memory-efficient fine-tuning through forward loss queries. In sparse mixture-of-experts (MoE) models, however, parameter perturbations can change the experts selected for a token. This path inconsistency mixes continuous parameter effects with discrete routing changes in finite-difference estimates. We propose ZOMoE, a ZO fine-tuning algorithm that combines expert replay with one-sided estimation. At each update, an unperturbed forward pass records the selected expert indices. Perturbed queries reuse these indices while recomputing routing weights and other continuous computations. Expert selections are refreshed at the next update, allowing routing to adapt throughout training. The recording pass also supplies a shared loss baseline, so replay adds no forward queries to the one-sided estimator. Mathematically, we isolate the term introduced by expert switching in the finite-difference estimate and show that expert replay removes this term. Across two sparse MoE backbones and six downstream tasks, ZOMoE achieves the highest average accuracy among the evaluated ZO methods, accelerates loss descent, and adds negligible memory or runtime overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.