Predictive Activation Steering via Solution Mapping and Stochastic Spectral Expansion
Abstract
Activation steering is a promising inference-time intervention that modulates large language models (LLMs) by perturbing their internal activations during generation. However, its practical application requires determining whether an effective steering intervention can be identified and quantifying trade-offs among suitable interventions. In this work, we propose an inexpensive framework for identifying optimized steering interventions subject to specified requirements. We model the steering coefficients as Gaussian random variables and approximate their effects on measured LLM behavior using a second-order Hermite polynomial expansion. Then, optimization is applied to find steering coefficients predicted to jointly satisfy specified targets. The resulting covariance matrix reveals structure within the steering coefficients, including coupling between coefficients and the degree to which they are jointly constrained. We propose a feasibility test that identifies unattainable specifications and systematically determines which requirements must be removed. We demonstrate the benefits of our approach through a case study and show that it performs competitively with existing methods for precise behavioral control of LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.