acceptodds
Under review as a conference paper at ICLR 2027

Concept J-Lens: From Token Binding to Concept-Level Control in the Jacobian Lens

Abstract

Chain-of-thought traces provide a useful but potentially fragile window into model reasoning, motivating complementary methods that read hidden states directly. The Jacobian Lens (J-Lens) maps hidden states to vocabulary-indexed directions that can be read and causally manipulated. Yet it remains unclear whether these directions capture the semantics underlying their indexing tokens or merely the tokens themselves. We test this distinction using writing interventions in Entity and Relation diagnostics, which require directions to generalize across alternative names and input-dependent relational outputs. Instead, J-Lens directions behave as token-level controls, predominantly promoting their construction tokens—a phenomenon we call token binding. Our analysis suggests that this behavior arises from fixed vocabulary targets and corpus-averaged Jacobians, neither conditioned on task examples. We therefore introduce Concept J-Lens (CJL), which maps task-defined targets through prompt-specific Jacobians and aggregates the resulting local directions into a reusable concept-level steering direction. Target–Jacobian ablations support this construction-level explanation. Across six language models and 25 tasks spanning five families, CJL raises average concept-level success from 4.18% to 60.47%, performs comparably to DiffMean activation steering overall, and is less sensitive to steering strength on Relation tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.