High-Dimensional Random Projection for Activation Steering in Language Models
Abstract
Activation steering has emerged as a key methodology for controlling the behavior of large language models (LLMs). Existing difference-in-means based methods, however, are fundamentally limited: they capture only mean differences between class activations and fail to recover discriminative signals that naturally exist in the nonlinear feature subspace under the superposition hypothesis. Rather than proposing another steering rule, we change the space in which existing rules operate. We propose **Hi**gh-**D**imensional **R**andom-projection for **A**ctivation Steering (HiDRA), a training-free plug-in that lifts activations into a higher-dimensional space through a random projection followed by an invertible nonlinearity, runs a base steering method's direction estimation and intervention in that space, and maps the steered activations back, leaving the base method's layers, token positions, and schedule unchanged. In the lifted space, difference-in-means directions provably capture discriminative structure beyond the reach of linear methods. Wrapping sequential steering (Mean-AcT), single-layer activation addition, and contrastive activation addition (CAA) with HiDRA strengthens behavioral control on jailbreaking, truthfulness, and multiple-choice behavioral benchmarks across the Gemma, Llama, and Qwen model families, with modest computational overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.