RouteTrace: Membership Inference from Routed Update Traces in Mixture-of-Experts Models
Abstract
Router updates in a sparse mixture-of-experts (MoE) model can reveal whether a record was used for post-training. We introduce ROUTETRACE, a white-box membership audit of paired base and post-trained checkpoints. ROUTETRACE measures how a candidate’s negative router gradient aligns with the net router update, both across layers and along the candidate’s route. The candidate’s token routes identify router rows associated with selected experts, allowing this alignment to be compared on selected and unselected rows. With three matched shadow checkpoints for each model and dataset, its learned scorers reach mean AUCs of .948 on Natural Questions and .862 on Dolly across Qwen, DeepSeek-MoE, and OLMoE. On Natural Questions, global depth features, which pool evidence across all router rows, preserve most of the full scorer’s ranking accuracy, while alignment on selected rows separates members more strongly than equal-sized next-ranked and random unselected controls
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.