acceptodds
Under review as a conference paper at ICLR 2027

Dense Credit Assignment for Adaptive MoE Routing via Layer-wise MoE-

Abstract

MoE LLMs are widely used to scale capacity under limited inference compute. Their activated expert count is a primary lever for that cost. Because expert-count demand varies across tokens and layers, static Top- overallocates routed compute. Recent work therefore studies token- and layer-wise count allocation. Among these, Ada-K (Yue et al., 2025a) trains a thin -selector with reinforcement learning from next-token probability. That objective optimizes generation quality via a terminal likelihood signal; it does not directly score the count decision at each (token, layer). We argue that denser MoE-latent credit better matches optimization, and propose -Credit: a reward that scores each layer's count decision by how closely the resulting MoE-block update (the change in hidden state across that layer's expert computation) matches a full-capacity (Top-) teacher. Unlike Ada-K's single terminal likelihood signal broadcast to every depth, this credit is a dense per-layer score along the student trajectory. With this training target on Gemma 4 26B-A4B (reasoning-mode decoding, 64k context), -Credit outperforms Ada-K on GPQA Diamond / MMLU-Pro / MMMU-Pro at a lower mean activated count (82.8/83.7/63.1 vs. 76.3/82.0/61.7; vs. ), reaching 98.8%/100%/94.7% of BF16 Top-8 accuracy at 70% routed MoE cost. The same recipe on Qwen3-30B-A3B (Qwen Team, 2025) (text-only) is again ahead of Ada-K on both GPQA Diamond and MMLU-Pro at comparable routed cost. These results support MoE- as an effective action-sensitive credit signal for learning dynamic under our evaluation protocol, and show that when the training signal aligns with count allocation, adaptive routing can cut activated experts more aggressively while better preserving downstream accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.