Direct Spectral Actuation: Prediction-Preserving Expert Differentiation in Mixture-of-Experts
Abstract
Trained mixture-of-experts (MoE) models may lack task-relevant expert differences, leaving routers with little useful gradient signal. Direct Spectral Actuation (DSA) adds protected expert-output residuals after training without changing the anchor prediction. Under a fixed function-space budget, DSA selects the write that maximizes added router-gradient power under the original task loss. This yields a leading singular mode at a symmetric anchor and a metric-aware trust-region solution at a general trained anchor. DSA freezes the residual adapter and updates only the reader; native experts and the pretrained router remain fixed. Controlled tasks and CIFAR-10-C test protection and the subsequent read, while an OpenML-CC18 replay tests selective intervention. On one frozen native top-8 OLMoE checkpoint, DSA reduces held-out NLL relative to the anchor by 12.4–25.8% on SST-2 and 40.7% on average on RTE. It improves on matched reader-only adaptation in all six evaluations and yields lower task-mean NLL than a matched final-layer RoMA-objective baseline on both tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.