acceptodds
Under review as a conference paper at ICLR 2027

Stabilizing Zeroth-Order Fine-Tuning of Mixture-of-Experts via One-Sided Estimation and Expert Replay

Abstract

Zeroth-order (ZO) optimization enables memory-efficient fine-tuning through forward loss queries. In sparse mixture-of-experts (MoE) models, however, parameter perturbations can change the experts selected for a token. This path inconsistency mixes continuous parameter effects with discrete routing changes in finite-difference estimates. We propose ZOMoE, a ZO fine-tuning algorithm that combines expert replay with one-sided estimation. At each update, an unperturbed forward pass records the selected expert indices. Perturbed queries reuse these indices while recomputing routing weights and other continuous computations. Expert selections are refreshed at the next update, allowing routing to adapt throughout training. The recording pass also supplies a shared loss baseline, so replay adds no forward queries to the one-sided estimator. Mathematically, we isolate the term introduced by expert switching in the finite-difference estimate and show that expert replay removes this term. Across two sparse MoE backbones and six downstream tasks, ZOMoE achieves the highest average accuracy among the evaluated ZO methods, accelerates loss descent, and adds negligible memory or runtime overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.