acceptodds
Under review as a conference paper at ICLR 2027

Steering Does Not Imply Localization: Multilingual Refusal Control in Mixture-of-Experts Models

Abstract

Intervening on a small set of experts in a mixture-of-experts (MoE) model can alter refusal behavior, but behavioral control alone does not show that refusal during unmodified inference depends on those experts. This distinction is especially important across languages: an intervention that works across the evaluated languages need not reveal a shared refusal mechanism. We construct an evaluation set by aligning requests from 11 benchmark collections across 12 languages and study Qwen3 and Gemma-4, two open-weight MoE models with manipulable sparse routing. Using multilingual routing contrasts, we select model-specific expert sets and fix them before evaluation. On 1,977 held-out English sources, forcing these experts increases observed refusal by about seven percentage points in both models and exceeds layer-matched controls. The original routers also select these experts during unmodified inference, yet removing them changes refusal much less than forcing does, with different levels of removal sensitivity in the two models. For corresponding requests in 11 non-English languages, forcing the selected experts likewise outperforms matched controls. The intervention also increases refusal of benign requests and should therefore not be interpreted as a net safety benefit. These results identify a multilingual interface for controlling refusal, but do not establish that the selected experts form a localized mechanism on which refusal depends.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.