acceptodds
Under review as a conference paper at ICLR 2027

Topology Before Steering: Localising and Gating Cyclic Concepts with Persistent Homology

Abstract

As Large Language Models become increasingly widespread, the need for effective safeguards and methods for controlling model behavior has grown rapidly. One prominent approach is activation steering, which directly intervenes in a model’s residual stream by adding a steering vector to encourage or suppress a target concept. Yet two fundamental questions arise before steering and have received comparatively little attention: where does the target concept reside within the prompt, and should we intervene on a given input at all? We study these questions for cyclic concepts, whose activations trace loop-like structures in representation space. We show that tools from Topological Data Analysis, particularly persistent homology, provide a principled way to address both problems. By analyzing activations across prompts, persistent homology identifies where a cyclic concept emerges and which tokens carry it. For new inputs, the same topological signal then determines whether the concept is present and whether steering should be activated, allowing the intervention to remain off when it is unnecessary. This yields a topology-aware approach to activation steering that jointly addresses both where and when to steer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.