acceptodds
Under review as a conference paper at ICLR 2027

A Manifold-Aware Activation Steering Method for Large Language Models via Geometric Safety Control

Abstract

Inference-time activation steering enables low-cost, plug-and-play behavior control by directly modifying the hidden states of large language models (LLMs). Existing methods steer hidden states in the high-dimensional Euclidean space, making them vulnerable to task-irrelevant perturbations. They also rely on overly strong and directionally coarse single-step vectors that can push trajectories away from semantically dense regions, impairing representation acquisition and generation quality. To address these issues, this paper proposes GSCSteer (Geometric Safety Control Steering), a manifold-aware activation steering method grounded in geometric safety control. GSCSteer formulates single-layer activation intervention as a control barrier function-based quadratic program (CBF-QP) over a low-dimensional semantic manifold. To suppress task-irrelevant perturbations, the Kernel-Based Semantic Control Barrier Function (KS-CBF) models a nonlinear semantic boundary, while Orthogonal Tangent Subspace Projection (OTSP) confines the induced state-dependent corrections to a low-rank proxy tangent space. GSCSteer then replaces coarse, overly strong one-shot updates with multi-step integration, using Retraction-like Radial Correction (RRC) to limit norm drift and produce finer trajectories that more closely follow the empirical semantic geometry. Experiments on four open-source LLMs show that, compared with existing activation steering methods, GSCSteer improves the win rate on helpfulness tasks by 1.5%–3.6%, improves on truthfulness tasks by 2.9%–16.0% while reducing perplexity by 22.1%–44.0%, and reduces toxicity on detoxification tasks by 2.7%–24.1%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.