The Second Bill: Pricing Activation Steering in the Monitor's Currency
Abstract
Activation steering is priced in one currency: how far an intervention moves the model's output distribution. Every linear activation monitor reading the same residual stream charges a second bill, poorly tracked by output-side quantities and computable in advance. The monitor covector pulls a probe's readout back to the intervention site at one vector–Jacobian product per input: its second moment is a metric on interventions, which a Gaussian model of a threshold frozen at a fixed false-positive rate turns into two break conditions. A monitor can be blinded or made to false-alarm, and, at a given displacement, its margin protects only against the first: the second happens at a displacement of 0.970 within-class spreads whatever the margin. On 666 probes trained to graded margins at up to four depths above every block, the false-alarm rate well past that displacement is 79%–100% in every margin bin, while the blinding rate falls from 99% to 6%–7%; with block fixed effects the coefficient on is for blinding and for false alarms. A standard difference-of-means vector, at the largest per-cell norm keeping generic KL and benchmark-accuracy drop within 0.05, makes 360 of 1,998 readings false-alarm, 22 of 618 on probes with ; the harmfulness probe crosses one norm further out, at 1.46× that capability budget. A covector-only screen catches 66% of those breaks at precision 88%. The joint budget pays 0.367× the output-optimal direction's monitor bill at 1.01× its KL, and on an unpriced monitor 0.579× Euclidean's, no better than a random-probe prior. Both break conditions and the in-budget false alarms replicate on three further models (1B–14B, pooled). All code and results will be released at https://github.com/xxx/xxx upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.