acceptodds
Under review as a conference paper at ICLR 2027

Native Routing Priors for Budgeted Low-Rank Adaptation of Mixture-of-Experts Models

Abstract

Under a low-rank parameter budget, adapting a mixture-of-experts (MoE) model couples expert coverage, local rank and expert identity. We examine these decisions through Native Router-Guided Layer–Expert Capacity Allocation (NRG), using reference-answer-free profiles and static DoRA adapters. On Qwen3.5-35B-A3B-Base, selected rank-32 adapters exceed approximately budget-matched all-expert adapters by 0.83–0.90 percentage points across three low-data VCR and NLVR2 settings. Grouped intervals support a positive conditional difference on NLVR2, while VCR intervals cross zero. Follow-up controls reveal that expert selection alone does not ensure favorable accuracy: on VCR Train50, selected rank-7 adapters fall 0.90 points below all-expert rank-7 adapters, while increasing rank on the same selected experts yields a 1.77-point gain. The corresponding NLVR2 rank comparison remains uncertain, and neither task establishes an advantage for routing-derived over uniform layer quotas. Historical execution compatibility is only partially verified, limiting attribution in these follow-up comparisons. Repeated random allocations retain a favorable VCR mean difference, but an exclusion-based probe does not identify a ranking mechanism. These observations motivate separate tests of capacity distribution and expert identity when evaluating routing priors. We provide budget accounting and numerical reproduction; the evidence remains conditional on observed models from one backbone and reused development subsets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.