acceptodds
Under review as a conference paper at ICLR 2027

Activation Flow: Manufacturing Activations for Steering

Abstract

Difference-in-means steering requires activations recorded while a model shows the desired behavior, which a sandbagging model withholds by deliberately underperforming. We introduce Activation Flow (ActFlow), which manufactures these activations from correct labels without fine-tuning. ActFlow sets target logits that rank each labeled item's correct answer first, and moves the logits toward them by adding one vector to all residual streams at one layer. ActFlow is a family of ordinary differential equations for , one for each rule that maps the required logit change to the velocity of . The smallest-norm rule lands exactly on the targets, while the others keep only the top singular directions of the Jacobian. We test ActFlow on three instruction-tuned models, each locked by a sandbagging prompt and by a password-locked LoRA. At , ActFlow keeping five singular directions raises the mean held-out ARC-Easy accuracy over the six locked models from to , against for fine-tuning and for the honest models. Furthermore, it scores higher than the smallest-norm rule in 16 of the 18 combinations of locked model and , and its steering direction is nearly orthogonal to the honest difference-in-means direction. It also unlocks two LoRA locks where the honest direction fails.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.