acceptodds
Under review as a conference paper at ICLR 2027

Calibrating Activation Steering with Sparse Behavioral Feedback

Abstract

Activation steering can control a frozen language model through an additive direction, yet a coefficient that works for one checkpoint can produce little change in another. We study how to turn this model dependence into a calibration problem. Across 15 instruction-tuned models and five personality attributes, residual-stream scale predicts sensitivity to a shared coefficient grid: a mid-layer RMS feature achieves a held-out-family rank correlation of 0.83. Within each model, this scale is stable across attributes, suggesting a reusable measurement for identifying candidate dose ranges. Behavioral feedback then selects a directional operating point within a range. In a retrospective evaluation of measured grids on five models, six rated responses from one calibration scenario select a dose that transfers to the remaining scenarios. At the same feedback budget, expanding the candidate grid raises the mean held-out signed contrast from 0.24 on the original grid to 1.27, with improvements in all 25 model–attribute averages. Six ratings retain 89.6% of the effect obtained using 30 ratings on the same scenario. These results separate model-level sensitivity from attribute-level dose selection and establish sparse behavioral calibration as a practical way to recover useful steering from an existing direction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.