CoFFiE: Consensus-First Fine-Tuning for MoE
Abstract
Mixture-of-Experts (MoE) models are increasingly used as base models for downstream tasks, but how best to fine-tune them remains an open question. Full fine-tuning stores optimiser state for every expert, even though only a few are active for each token and each expert sees only a small fraction of the fine-tuning data. Parameter-efficient methods reduce this cost, but most inherit designs from dense models: each expert or adapter computes its update independently, whereas the output of an MoE layer depends on how its active experts work together. We propose CoFFiE (Consensus-First Fine-tuning for Mixture-of-Experts), which keeps the experts and the router frozen and adapts this coordination instead. A small shared network refines each expert's output using its own state and the route-weighted consensus of the active experts. Because refining the experts also changes this consensus, we define the refined outputs implicitly, as their joint fixed point. A few iterations suffice in practice, and training memory does not grow with the iteration count. We further extend the method by propagating consensus information across layers, which also improves accuracy significantly. On OLMoE-1B-7B, Mixtral-8x7B and Qwen3-30B-A3B, CoFFiE outperforms the strongest MoE-specific PEFT baselines with the same or fewer trainable parameters, by up to 3.3 points on commonsense and arithmetic reasoning benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.