acceptodds
Under review as a conference paper at ICLR 2027

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

Abstract

Controlling the behavior of large language models (LLMs) at inference time is central to alignment, but existing steering methods such as Contrastive Activation Addition (CAA) rely on fixed single-layer interventions derived from aggregate activation differences. A single direction is applied to semantically diverse inputs, and the induced change is not sustained as representations propagate through later layers, which limits both the strength and the reliability of the resulting steering. We introduce CircuitSteer, a framework that uses Sparse Autoencoders (SAEs) to identify and manipulate coherent semantic circuits distributed across multiple layers of the residual stream. We construct an aligned co-activation circuit from feature co-activation and the geometric alignment of decoder directions, isolating the multi-layer subcircuits that jointly mediate a given target behavior. We then synthesize dense steering vectors from these sparse features and apply coordinated multi-point interventions along the circuit, guiding the model's internal semantic trajectory rather than perturbing a single site. We evaluate CircuitSteer with contrastive examples on three distinct behaviors, toxicity, emotion intensity, and sycophancy, using Gemma-2-2B and Llama-3.1-8B-Instruct. CircuitSteer is the only method that significantly reduces the target behavior while preserving fluency in every configuration; each baseline either degrades text quality or fails to shift the behavior on at least one configuration. These results show that multi-layer circuit steering, enabled by enforcing geometric alignment among selected features, yields more robust and more consistent behavioral control than static single-point interventions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.