acceptodds
Under review as a conference paper at ICLR 2027

Robust Steerability: Closed-Loop Control of Language Model Representations Under Distribution Shift

Abstract

Activation steering can change the behavior of large language models, but an intervention calibrated on one prompt distribution may fail when deployment inputs alter the network's internal dynamics. Neurostimulation faces an analogous problem and has increasingly moved from fixed, open-loop stimulation toward closed-loop, state-dependent control. Following this principle, we formulate activation steering as finite-horizon feedback control over transformer depth. We identify local layer-to-layer dynamics from calibration activations and represent their mismatch with the true network through a disturbance channel fitted to the empirical covariance of one-step prediction residuals. We then synthesize an controller that minimizes the worst-case amplification of this mismatch, and define robust steerability, , as the inverse of the smallest achievable amplification. We show that the associated certificate localizes fragility to specific layers and disturbance directions, and that nominal LQR steering is the large- H_\infty steering reduces mean jailbreak attack success from 9.3\% to 1.7\% relative to the strongest baseline, improves truthfulness\(\times\)informativeness under Spanish-prompt transfer by 5.4 points on average, and retains 77--95\% of unsteered math accuracy under language steering, compared with 0--85\% for A-LQR. Gains under context-length shift are model-dependent. Finally, S_rob $ varies systematically with the steering task and increases with model scale within a task. These results suggest that modeling uncertainty matters most when steering is applied outside its calibration regime, and that biological and artificial neural networks pose related control problems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.