How Far Can Instruction Steering Go? Coverage, Quality, and Calibration
Abstract
Activation steering promises modular control of language models: learn one direction per instruction, then add the directions that a request needs. This promise rests on two assumptions: directions that work alone also work together, and an answer that passes an instruction checker is still a good answer. We test both assumptions on seven models. We calibrate one edit per instruction and freeze it. An answer counts as a success only if it passes every checker and a blind quality judge, which agrees with a human audit. We separate installation, where the edit must carry an instruction that the prompt omits, from reinforcement, where the prompt already states it. Few instructions can be installed reliably. The quality judge shrinks this set further. On four smaller models, it removes all four installations that the automatic checks accept, and it roughly halves the usable reinforcements. Where installations do succeed alone, as on Gemma-3-27B, they also compose. On instruction pairs, they beat the unedited model by 35.0 points. Still, they trail simply stating the instructions by 17.2 points. Reinforcement gives mixed results and falls behind prompting as sets grow, yet much of this loss comes from calibration. Composing steering directions is therefore feasible, but its value depends on coverage, answer quality, and calibration, and prompting remains the baseline to beat.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.