SentinelMark-MoE: Auditing Mixture-of-Experts Route Execution with Secret Challenges
Abstract
Mixture-of-experts (MoE) models route each token to only a small set of experts. This saves computation, but it separates the route chosen by a trusted router from the work carried out by a distributed backend. A record of the intended route cannot reveal whether the backend rerouted, omitted, or replaced expert work. Existing model watermarks mainly test ownership or generated text, not the route that a serving system actually executes. We introduce SENTINELMARK-MoE, an online audit for MoE route execution. For each decision, it forms two routes that satisfy the routing constraints and uses secret scores under a fresh tag to choose one. A trusted observer records a small summary of the route that ran. The detector checks this summary for the expected secret pattern. It can raise an alarm at any time while controlling false alarms. It skips pairs with identical summaries and uses the score distribution to better detect infrequent changes. With fresh random scores, selection preserves the chosen routing distribution exactly; a pseudorandom function provides a computational version of this guarantee. Experiments across three MoE families and five language benchmarks show that SENTINELMARK-MoE detects changes in expert execution while maintaining model quality. It consistently outperforms the tested statistical detectors against rerouting, omitted expert work, and expert-identity substitution. The results high- light two practical findings: model-specific calibration keeps routing changes small, and the choice of trusted observation determines which attacks the audi- tor can detect. Matched-prompt evaluations further show the trade-offs between detection, model quality, and serving cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.