acceptodds
Under review as a conference paper at ICLR 2027

HeteroLoRA-MoE: Heterogeneous Low-Rank Experts with Diverse Projection Patterns and Efficient GPU Kernels

Abstract

Expert differences in a pretrained mixture-of-experts (MoE) model exhibit low-rank structure, motivating compact expert representations based on shared weights and low-rank updates. Motivated further by variation in this structure across projections and experts, we present HeteroLoRA-MoE, which distributes expert-specific low-rank updates across four balanced projection patterns over shared dense SwiGLU weights. A single token-level router selects complete experts and combines their outputs. This compact design introduces small matrix products and projection-dependent scheduling, so parameter savings do not automatically translate into lower latency. To address these execution costs, we develop HeteroKernel, a structure-aware GPU implementation that reuses shared gate and up results and aggregates expert intermediates before applying the shared dense down projection. Projection-aware grouping, fused reductions, and pipelined weight loading organize the remaining adapter computation, skipping absent branches while preserving the mixture algebraically. In four-seed mixed-SFT experiments initialized from Qwen2.5-0.5B/1.5B-Instruct, HeteroLoRA-MoE achieves higher mean scores than a parameter-matched MixLoRA-style baseline on multiple evaluation tasks at both model scales. On NVIDIA L20, HeteroKernel reduces the overhead of routed low-rank computation, achieving 1.85–6.79x speedups over the PyTorch implementation across the tested prefill and one-token decode settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.