Steering Vectors Can Control What Models Learn in Supervised Fine-Tuning: An Analysis of Behavior Offloading
Abstract
Fine-tuning and activation steering provide two distinct ways for a model to express a behavior. Fine-tuning writes behavior into the model's parameters, while steering vectors can supply behavior through the residual stream without changing the parameters. We study what happens when these two mechanisms overlap during training. If a behavior required by the training targets is already supplied through the activations, how much of it is still learned into the weights? We call the reduction in learned behavior under this intervention behavior offloading. Across eight behaviors spanning format markers, code constructs, style, toxicity, and reasoning, and across models from three families up to 14B parameters, we find that behavior offloading extends far beyond the coarse persona traits it was first shown on and, when successful, reduces the behavior further than a random direction of the same norm does while largely preserving task performance. We then characterize when this control remains effective and where it breaks down. A direction that steers only over a narrow magnitude range can fail to offload even when a more stable construction of the same behavior succeeds. Among successful offloads, the onset occurs within a relatively narrow raw-magnitude range within a model, while the usable range before coherence breaks varies substantially across behaviors. Inference-time steering strength is also associated with whether a direction supports clean offloading or suppresses the behavior only after task performance degrades. Finally, steering does not create a behavior that is absent from the training targets in any setup we test. Together, these results show that activation steering can control which supervised behaviors are written into the weights, but does not replace supervision Sample code is available at https://anonymous.4open.science/r/Behavior-Offloading-FB6F/README.md..
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.