acceptodds
Under review as a conference paper at ICLR 2027

Isotropic Steering of Large Language Models

Abstract

Activation steering modifies intermediate model activations during inference, providing a promising solution to intervene in the behavior of large language models (LLMs) without updating model parameters. However, interventions from existing steering methods often suffer from strong anisotropy, which harms steering robustness and transferability. To address this issue, we propose an Isotropic Steering (TRIS) method for LLMs that calibrates intervention strengths across different steering directions and different activations. In particular, TRIS preconditions the gradient of a steering objective with regularized covariance of activations and obtains normalized interventions in coordinates aligned with the activation distribution. From an optimization perspective, we show that TRIS is equivalent to Steepest Ascent under the Mahalanobis metric. Across four LLMs from the Mistral, Llama, and Qwen families, TRIS improves the primary task metric over ODESteer in 10 of 12 model–task comparisons. In addition, we find that covariance estimated on one task transfers effectively to other tasks within the same model, suggesting that TRIS captures reusable activation distribution and transfers well across tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.