acceptodds
Under review as a conference paper at ICLR 2027

From Capability to Alignment: Repurposing MoE Routing for Safety

Abstract

Mixture-of-Experts (MoE) models route each input to a sparse subset of experts, re- alizing capability through per-input expert specialization. Can this input-dependent routing also be used to improve safety? On a given input, some experts contribute to unsafe behavior while others support safe responses, suggesting that suppress- ing the former could improve safety. Realizing this idea raises two challenges. First, the identity of risk-amplifying experts shifts across queries, so the experts to suppress must be identified per input. Second, queries vary in risk level, so suppression strength must also be calibrated per input. Risky inputs may require stronger intervention, while benign ones should remain largely unaffected. To address these challenges, we introduce PIRA (Per-Input Router Alignment), an alignment-enhancing technique that adapts both the experts suppressed and the sup- pression strength at inference time, without updating model weights. PIRA derives two complementary signals from the same input representation: gradients of an output-safety score with respect to router logits identify active experts that amplify risk for the current input, while an input-risk score controls intervention strength. For queries above a calibrated risk threshold, a single forward-backward probe computes a per-query negative router bias that remains fixed during generation; other queries receive no suppression. Across three MoE backbones and six selected PandaGuard JailbreakBench attacks, PIRA improves average safe-response rates by over 20 percentage points relative to the original models. XSTest answer rates decrease by at most 2 percentage points, and MMLU accuracy by at most 0.3 percentage points. Additional evaluations further support the generality of PIRA across model backbones and safety settings. These results suggest that MoE’s per-input routing mechanism can be repurposed to enhance safety.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.