JORA: JOINT EXPERT-OUTPUT REWEIGHTING FOR TRAINING-FREE MOE ACCELERATION
Abstract
Skipping experts lowers the cost of mixture-of-experts (MoE) inference, but can substantially degrade model quality. The resulting approximation depends on the retained experts and the coefficients used to combine their outputs. We intro- duce JORA, a training-free method that jointly estimates these coefficients after expert selection to approximate the original routed mixture. Offline calibration records expert-output inner products. At inference, these statistics and the cur- rent routing weights define a small regularized system, avoiding evaluation of omitted experts. Model parameters remain unchanged. JORA supports per-token prefill retention and shared-pool decoding; our analysis establishes optimality for the fixed-set calibration objective and characterizes sensitivity to calibration mis- match. Across text and vision-language benchmarks, JORA outperforms the eval- uated expert-skipping and re-routing baselines in most task–budget comparisons, with improvements under prefill-only, decode-only, and joint reduction. In single- GPU vLLM serving, Qwen3-30B-A3B with native prefill and a 32-expert decode pool achieves a maximum observed end-to-end speedup of 1.31×over Native for 256 input and 1,024 output tokens at concurrency 64.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.