acceptodds
Under review as a conference paper at ICLR 2027

Endpoint-Sensitive Causal Audits of Mixture-of-Experts Safety

Abstract

Expert-routing interventions in autoregressive mixture-of-experts models can alter both the next-token prediction and the attention state retained for later decoding, while the measured safety effect further depends on whether a prefix or the complete answer is evaluated. We introduce a four-cell token/state experiment that isolates these pathways, using prefix and complete-answer endpoints, paired and different-request donors, and a final-layer continuation invariant. In a three-model development audit with 279 prompt pairs per model, changing the evaluated span alters which pairs appear unstable even when aggregate flip counts agree. A frozen 49-request study then shows that, with the next token fixed, paired routing transfer changes decoded text in 32/33 naturally completed OLMoE misinformation comparisons and 20/21 Qwen3-MoE comparisons, yet changes no complete-answer compliance labels; a dense residual comparator similarly alters text in 19/31 comparisons without affecting compliance, and different-request donors produce comparable text changes. Few native MoE answers are classifier-positive, limiting conclusions about repair, while one reserved OLMoE case exhibits a nonzero state contrast under prefix grading and zero under complete-answer grading (with a selected development case showing the complementary pattern), and execution audits identify four auxiliary factuality trials that fail strict position isolation. These findings provide bounded counterexamples to treating token recovery, changed text, and measured safety repair as interchangeable evidence, and support joint audits of decoding pathways, numerical validity, and evaluation endpoints.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.