acceptodds
Under review as a conference paper at ICLR 2027

Jᴿ-Space: Recursive Refinement and Interpretation of Language Model Steering

Abstract

Activation steering modifies language model behavior through additive interventions, but a fixed steering vector may leave room for further improvement as the intervention changes the model’s internal state. We introduce \(J^R\)-Space, a framework for recursively refining existing steering vectors and examining their corrective directions. At each calibration round, \(J^R\)-Space recomputes behavioral gradients under the current intervention, projects them away from previously used directions, and constructs a sparse signed correction from a vocabulary-indexed Jacobian dictionary. The resulting vectors remain fixed at deployment and require no test-time gradients. Experiments on two model families across 50 behavioral tasks and full-answer TruthfulQA show improvements across multiple starting methods, with ablations examining gradient recomputation, update staging, and history projection. Building on the same dictionary, we develop a question-level diagnostic that ranks token directions by their alignment with corrective gradients. Across seven behavioral tasks, this analysis reveals task-related lexical structure, with individual cases showing agreement between gradient-match signs and target-relative meanings. Together, these results demonstrate effective recursive refinement and provide a vocabulary-based view of the directions that can further influence steered behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.