ZOSteer: Steering Language Models from Scalar Rewards Alone
Abstract
Activation steering controls a language model by adding a fixed vector to its residual stream, but existing methods derive this vector from contrastive prompt pairs or differentiable objectives, leaving behaviors specified only by a scalar, non-differentiable reward largely out of reach. We introduce ZOSteer, which estimates a steering vector as the zeroth-order gradient of the reward in activation space. We inject isotropic Gaussian perturbations at a single layer, score each perturbed model with the reward, and take the cross-covariance between perturbations and rewards. This estimates the gradient of the Gaussian-smoothed reward from forward passes alone, without contrastive data or backpropagation. Because the reward is an average over items, the same perturbations also yield a gradient for every training item at no extra cost. The residual spectrum of these item gradients reveals sub-populations whose gradients conflict with the majority, and reweighting them repairs the pooled direction without new forward passes. Across three 1B-scale models and four benchmarks, ZOSteer gives the largest gains on GSM8K for every model, and its reweighted variant closes most of the gap to the strongest baseline on bAbI QA20 and surpasses it on CUTE SubChar, typically with fewer perturbations than the hidden dimension. These results suggest that what a single vector can steer depends on how a task's items agree, not only on how the model represents the capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.