acceptodds
Under review as a conference paper at ICLR 2027

Introspecting Alignment Shifts Beyond Behaviors Implanted Through Fine-Tuning

Abstract

Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that were present in the source model. Recent work has shown that large language models can be trained using LoRA-based modules known as introspection adapters (IAs) to describe behavioral changes induced by fine-tuning. However, existing studies primarily consider settings in which the model is fine-tuned on datasets explicitly designed to implant a specific behavior and is then asked to explain the implanted behavior. This differs from practical deployment scenarios, where the misalignment is not necessarily deliberately implanted, but arises only as a side effect. To bridge this gap, we formulate a novel problem setting, in which the target of introspection is not necessarily a behavior explicitly implanted through fine-tuning, but rather alignment shifts that may emerge as unintended side effects, and we construct a dataset for this setting. Furthermore, to enhance sensitivity to internal model changes, we propose the Delta-Aware Introspection Adapter (DAIA), a novel mechanism designed to explicitly process both base-model activations and activation differences induced by fine-tuning. Our empirical evaluation shows that introspection learning generalizes to unseen fine-tuned models and safety categories, and that DAIA generally outperforms existing IAs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.