RouteMIA: Expert Routing Leaks Membership in MoE LLMs
Abstract
Many membership inference attacks on large language models rely on log-probabilities of the input tokens that deployed services do not expose. Mixture-of-Experts (MoE) models, however, leak a different MIA signal: how input tokens are distributed across experts, which may be exposed through side channels even when log-probabilities are hidden. Specifically, we find that fine-tuning shifts this distribution more for members than for non-members. Building on this observation, we propose RouteMIA, which infers membership by measuring the Jaccard distance between expert sets selected by a public base model and by its fine-tuned counterpart from the same input. Unlike prior attacks, RouteMIA uses only the number of tokens routed to each expert, a signal that side channels may reveal, rather than log-probabilities. Because some routing changes occur regardless of membership, RouteMIA calibrates these changes against token identity and base routing margin using non-member reference data. Experiments across three MoE LLMs and four datasets demonstrate routing-based membership leakage and show that the proposed calibration generally improves attack performance. Our findings identify expert routing as a new privacy attack surface introduced by sparse MoE architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.