acceptodds
Under review as a conference paper at ICLR 2027

RouteTrace: Membership Inference from Routed Update Traces in Mixture-of-Experts Models

Abstract

Router updates in a sparse mixture-of-experts (MoE) model can reveal whether a record was used for post-training. We introduce ROUTETRACE, a white-box membership audit of paired base and post-trained checkpoints. ROUTETRACE measures how a candidate’s negative router gradient aligns with the net router update, both across layers and along the candidate’s route. The candidate’s token routes identify router rows associated with selected experts, allowing this alignment to be compared on selected and unselected rows. With three matched shadow checkpoints for each model and dataset, its learned scorers reach mean AUCs of .948 on Natural Questions and .862 on Dolly across Qwen, DeepSeek-MoE, and OLMoE. On Natural Questions, global depth features, which pool evidence across all router rows, preserve most of the full scorer’s ranking accuracy, while alignment on selected rows separates members more strongly than equal-sized next-ranked and random unselected controls

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.