Interpretability Is Not Just a Tax: Sparsity and Diversity Yield Legible and Composable Reward Experts for Personalization
Abstract
Preference modeling aims to capture heterogeneous human values and increasingly relies on learning latent components to represent this variation without fine-grained supervision. What these components actually represent, however, is rarely examined. We audit Mixture-of-Experts (MoE) reward models along three rungs—semantic coherence, functional correspondence, and composability—in controlled and real-world settings, holding the architecture and training data fixed so that differences isolate the effect of the training objective. Our diagnosis is that the components conventional MoE reward models learn are entangled, which limits not only interpretability but also the personalization such models are built for. The entanglement is not intrinsic to the architecture but a property of the training objective: off-the-shelf sparsity and diversity regularizers already substantially reduce it. The resulting configuration, which we call sparse MoE, recovers all underlying semantic categories without domain supervision in a controlled setting, reaching 84.8% expert purity against 60.8% for conventional MoE; on real-world data its experts admit human-interpretable task- and domain-level descriptions aligned with their functional strengths. Targeted interventions on routing weights produce predictable preference changes, showing that these interpretations are behaviorally grounded rather than merely descriptive. Crucially, the disentanglement induced by this objective also makes the experts composable: lightweight router updates yield a 25.81-point improvement in few-shot personalization with 50 examples, roughly 2.7 times the gain of the identical architecture and adaptation procedure applied to entangled experts. Together, these results show that MoE reward models can learn interpretable expert structure, and that the same structure can substantially improve personalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.