acceptodds
Under review as a conference paper at ICLR 2027

Channel Selection for Mixture-of-Experts Decoding Needs No Calibration Data

Abstract

Mixture-of-experts (MoE) language models decode slowly because every generated token must read its experts' weights from memory. Yet for a given token, a small share of each expert's channels carries most of its output. Methods that skip the other channels must choose which to compute before reading them, and the strongest tune thresholds on calibration data. We show that these data are unnecessary: replacing the calibrated thresholds of the strongest such method, Prox, by a fixed number of channels per token (top-k) changes its divergence from the full model by at most a quarter, whereas changing how channels are ranked changes it more than sixfold. Thresholds also make the cost depend on the input: on Gemma 4, Prox's thresholds read fewer bytes on reasoning than on their calibration data and collapse into repetition, which top-k selection with the same score avoids at the same reads. We therefore build a reader that ranks channels from the weights alone with a compact ternary sketch (7.8% of the expert bytes), computes the top-ranked channels exactly and recovers the rest with a low-rank map of the weights. Across four MoE models and nine baselines at equal memory reads, it is closer to the full model than Prox from 18% of expert reads up on OLMoE and Gemma 4 and level on the Qwen models, reads the same bytes on every input, and speeds up the MoE block by up to 1.9x over optimized vLLM. Channel selection can thus be applied to a new checkpoint from its weights alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.