Beyond Independent Top-: Pairwise Energy Routing for Frozen Mixture-of-Experts
Abstract
Sparse mixture-of-experts (MoE) models activate only of experts per token, yet conventional Top- routing ranks experts individually without explicitly accounting for interactions within the selected subset. The highest-scoring experts therefore need not form the most effective subset. We propose PairQ, a parameter-efficient method for adapting pretrained MoEs through token-conditioned pairwise routing. PairQ combines frozen native-router scores with learned interactions between expert-selection variables, assigning a quadratic energy to each subset. Only the added routing parameters are trained; the backbone, experts, native router, and expert-mixing rule remain unchanged. An energy-weighted surrogate over candidate subsets propagates language-model loss gradients while preserving hard selection in the forward pass. The selection objective admits a quadratic unconstrained binary optimization (QUBO) formulation. Evaluations on Qwen3-30B-A3B and GPT-OSS-20B show gains over native routing across five mathematical reasoning benchmarks. A complete autoregressive inference run demonstrates that a physical coherent Ising machine (CIM) can decode the routing QUBO.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.