acceptodds
Under review as a conference paper at ICLR 2027

RcMoE: Semantic-Conditioned Retrieval for Efficient MoE Inference on Commodity Edge Devices

Abstract

Mixture-of-Experts (MoE)-based large language models (LLMs) leverage sparse activation to increase model capacity without proportionally increasing per-token computation. However, their large parameter footprint and storage demands hinder deployment on resource-constrained edge devices. Existing on-device MoE approaches remain largely computation-centric, limiting the efficiency benefits of sparse activation under tight edge resource constraints. Although lookup-based MoE methods can replace expert computation with precomputed expert outputs, existing token-indexed schemes limit retrieval expressiveness and shift the bottleneck to storage overhead, fragmented execution, and data movement. To address these challenges, we propose RcMoE, a retrieval-centric on-device MoE inference framework that constructs a semantic-conditioned lookup table (SC-LUT) of precomputed expert outputs, thereby improving retrieval expressiveness. RcMoE further develops a lightweight retrieval execution pipeline with SC-LUT compression, a fused retrieval kernel, and asynchronous prefetching to reduce storage footprint, execution fragmentation, and data movement overhead. Experiments on commodity edge devices demonstrate that RcMoE achieves up to 30.6 the throughput of DMoE while maintaining competitive accuracy under edge resource constraints.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.