RRD: Routing-and-Residual Distillation for Recovery after Dense-to-MoE Conversion
Abstract
Mixture-of-experts (MoE) models reduce per-token feed-forward network (FFN) computation by activating only a subset of experts. Dense-to-MoE conversion reuses pretrained dense weights, but sparse top- execution can create a gap between the converted student and its dense teacher. We formulate recovery under a fixed activation budget as two linked objectives: selecting useful experts and matching layerwise FFN outputs. Cross-entropy (CE) and logit-based knowledge distillation (logit-KD) supervise next-token predictions but do not directly constrain either objective. We propose Routing-and-Residual Distillation (RRD), which derives top- routing targets from teacher neuron activations and matches the realized sparse FFN output to the corresponding dense teacher FFN output. RRD combines these internal objectives with CE and logit-KD while retaining the converted expert partition and fixed activation budget. Across Qwen2.5-7B and Llama-2-7B, RRD achieves the highest mean zero-shot accuracy among the evaluated methods at both 25% and 50% FFN sparsity. On Qwen2.5-7B at 25% sparsity, RRD reaches 68.40% mean zero-shot accuracy after training on about 40M C4 tokens, outperforming the strongest same-topology baseline by 3.13 percentage points. At 50% sparsity, it outperforms the strongest compute-matched baseline by 2.47 percentage points under an approximately matched training-FLOP budget that includes the frozen teacher's forward pass. These results support direct supervision of expert selection and layerwise FFN outputs during sparse recovery. Code is available at https://anonymous.4open.science/r/routing_and_residual_distillation-7226/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.