A Time–Frequency Perspective on Activation Steering in Large Language Models
Abstract
Activation steering offers a lightweight way to control the behavior of large language models by modifying hidden states during inference. We investigate how changes in these states across token positions can guide such interventions. Taking a time–frequency perspective, we apply the discrete cosine transform (DCT) along token positions to separate the mean hidden state from variation at different frequencies. An analysis of 500 concepts from Concept500 reveals attribute information in nonzero-frequency coefficients and shows that this information is associated with patterns of hidden-state variation. These findings motivate SARA, a training-free method that uses recent hidden-state changes as feedback to adjust a base steering direction. During generation, it computes this feedback by comparing nonzero-frequency coefficients from recent hidden states with statistics estimated from the calibration set. Across three instruction-tuned models, SARA achieves the highest scores among the evaluated methods on all three TruthfulQA multiple-choice metrics. On the math and code subsets of Concept500, it also outperforms the trained baselines in concept relevance and aggregate score.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.