Steering Does Not Imply Localization: Multilingual Refusal Control in Mixture-of-Experts Models
Abstract
Intervening on a small set of experts in a mixture-of-experts (MoE) model can alter refusal behavior, but behavioral control alone does not show that refusal during unmodified inference depends on those experts. This distinction is especially important across languages: an intervention that works across the evaluated languages need not reveal a shared refusal mechanism. We construct an evaluation set by aligning requests from 11 benchmark collections across 12 languages and study Qwen3 and Gemma-4, two open-weight MoE models with manipulable sparse routing. Using multilingual routing contrasts, we select model-specific expert sets and fix them before evaluation. On 1,977 held-out English sources, forcing these experts increases observed refusal by about seven percentage points in both models and exceeds layer-matched controls. The original routers also select these experts during unmodified inference, yet removing them changes refusal much less than forcing does, with different levels of removal sensitivity in the two models. For corresponding requests in 11 non-English languages, forcing the selected experts likewise outperforms matched controls. The intervention also increases refusal of benign requests and should therefore not be interpreted as a net safety benefit. These results identify a multilingual interface for controlling refusal, but do not establish that the selected experts form a localized mechanism on which refusal depends.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.