Representational Inertia: Activation Stitching Reverses Fine-Tuned Behavior in Large Language Models
Abstract
Fine-tuning enables general-purpose language models to adapt to specific tasks and align with human preferences such as safety and truthfulness. Existing work has examined the representational changes induced by this process, yet it remains unclear how earlier representations remain functionally usable after training. We find that fine-tuned models retain backward compatibility with their precursor checkpoints, a property we call representational inertia. We investigate this property through cross-checkpoint activation stitching (CAST), which supplies precursor activations at a selected layer while retaining the fine-tuned computation downstream. Across six pairs of base and aligned models, a single stitch at intermediate depth returns the model to precursor behavior, at depths that vary across models and behaviors. The fine-tuned blocks that still run downstream cannot restore what training installed. The reverted text remains fluent, indicating that those blocks process the precursor's representation rather than being disrupted by it. This compatibility exposes a vulnerability in the activation space of aligned models, and it likewise allows accuracy suppressed by three unlearning methods to be recovered. It also enables possibilities to repair the models, as clean precursor activations suppress emergent misalignment and trained backdoors. Finally, the same intervention can be trained against, and we showcase that a model trained to recover from it also becomes more robust to unseen textual jailbreaks. Our findings show that fine-tuning installs behavior without displacing the computation it was trained on, with consequences for both the robustness of desired behavior and the correction of harmful changes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.