acceptodds
Under review as a conference paper at ICLR 2027

RouteVerify: Routing Interventions for Step-Level Verification in Mixture-of-Experts Models

Abstract

Routing states in mixture-of-experts (MoE) models can distinguish correct from erroneous reasoning steps, but such information does not by itself show whether changing routing changes a model's step-level verification judgment. We introduce RouteVerify, a controlled intervention study that replaces selected internal activations from one solution trace with activations from another, while keeping the evaluated input unchanged. Across nine MoE checkpoints, we compare three intervention interfaces: router logits, MoE outputs, and post-MoE residual states, and vary the layers and token positions at which activations are replaced. We measure changes in a verification margin that quantifies the model's preference for INCORRECT over CORRECT. We also evaluate the same input with activations from correct versus erroneous steps to test whether the resulting judgment depends on the source step's correctness. At the same layer and token positions, post-MoE residual replacements produce larger paired margin effects than router-logit or MoE-output replacements. Router-logit interventions become substantially stronger when extended across multiple routed layers and tokens, and intervening on equal numbers of tokens at different positions produces systematically different effects. However, large paired effects do not necessarily imply consistent differences between correct and erroneous source steps. With the evaluated input fixed, erroneous-step activations produce more INCORRECT-favoring margins than correct-step activations under local residual interventions in most checkpoints; this ordering is less consistent under broad routing interventions. These findings distinguish changing a verification judgment from changing it according to the source step's correctness, and show that both depend on the intervention interface and its layer and token coverage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.