Large Language Lobotomy: Locating and Silencing Safety in Mixture-of-Experts LLMs
Abstract
Mixture-of-Experts (MoE) language models route each token through only a small subset of experts, so the computation a prompt uses is conditional on its content and need not represent refusal uniformly. Prior attacks on MoE safety largely identify safety-relevant components from static activation frequencies; we ask instead which information in the routing trace best locates the components whose suppression changes refusal. We introduce Large Language Lobotomy (), a white-box inference-time framework that requires no target-model fine-tuning or parameter updates. records token-resolved routing trajectories for malicious prompts and minimally changed benign twins, trains a local-expert LSTM (LE-LSTM) to predict refusal, attributes the prediction back to individual local experts, and iteratively silences the highest-ranked ones. Across eight open-weight MoE LLMs, raises the average attack success rate (ASR) from % to %, reaching %, while accuracy on five downstream benchmarks falls by % on average at the most aggressive intervention we evaluate; it outperforms prior routing-based attacks that rely on static activation frequencies and matches F-SOUR, the strongest per-prompt routing search, but leverages one reusable expert ranking. With the silencing procedure held fixed, localization degrades when token order is destroyed and degrades further when routing is pooled across the prompt. The informative structure is the ordered path a prompt takes through the experts, which static usage counts discard. Token-resolved analyses further show that refusal evidence emerges at context-dependent positions before the end of the prompt. These results identify token-resolved routing trajectories as a strong signal for localizing intervention-relevant safety computation in MoE LLMs, and show that this signal suffices to jailbreak them without changing a single weight.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.