acceptodds
Under review as a conference paper at ICLR 2027

The Price of a Miss: Serving-Aware Expert Routing for Edge MoE Inference

Abstract

Expert-cache misses connect routing decisions to the cost of serving Mixture-of-Experts models under a GPU memory budget. We present TollRoute, a training-free routing rule that replaces a missing expert with a resident expert only whenits router logit lies within a fixed margin of the best desired expert. The rule guar-antees conditional miss non-increase and retains at least an e−γ fraction of thedisplaced expert’s routing weight for each substitution. Its serving interface useslayer–expert cache keys, admits only experts transferred to the GPU, and cali-brates the transfer/CPU split with the target platform’s expert-processing paths.The evaluation pairs per-layer and global LRU caches at equal byte budgets onGranite-3.1-1B-A400M, with a single NVIDIA B300 and sequential teacher forc-ing. At γ = 0.5, the per-layer comparison combines a mean CE increase of atmost 2.0 mnats with relative miss reductions of 10.0–19.3%. A serving case studyconnects these routing statistics to transfer, CPU execution, and routing overhead,while retaining the integer allocation of missing experts. The resulting analysisidentifies when fewer misses reduce the exposed service cost and when eliminat-ing the final miss in a layer is decisive.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.