acceptodds
Under review as a conference paper at ICLR 2027

WaXup: Weight Averaging of Experts as Data Augmentation for Robust Introspection

Abstract

As the capabilities of large language models keep improving, traditional mechanistic interpretability methods struggle to keep pace. In contrast, trainable introspection methods emerge as a promising alternative; they repurpose the model itself as a self-diagnostic tool, trained to verbalize its internal mechanisms into natural language explanations. While training on datasets of finetuned experts achieves encouraging results, it currently lacks robustness and generalization. Building on the data augmentation framework of MiXup, we introduce WaXup, a method that augments the introspection training set through linear interpolation in the weight space. This is motivated by the linear mode connectivity property: weight averaged experts interpolate the behaviors of the initial experts. We evaluate WaXup on a hidden-behavior benchmark and a meta-dataset of real-world NLP tasks. WaXup improves the accuracy of generated descriptions as well as robustness to various distribution shifts. We hope to improve the reliability of trainable introspection methods towards better and safer models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.