Where You Steer Shapes the Direction: Operator Conditioned LLM Behavior Control and J-lens Interpretation
Abstract
Activation steering controls LLM behavior by modifying internal representations. However, existing methods primarily focus on identifying behaviorally relevant directions, while overlooking how intervention operators, such as steering during prefill versus decode, shape both the construction and control effectiveness of these directions. This gap motivates us to reformulate activation steering as a local control problem jointly defined by an objective and a specific intervention operator, proposing Operator-Conditioned Gradient Steering (OCGS). OCGS constructs a differentiable behavioral objective from high- and low-scoring target-model responses and naturally aggregates its state gradients over operator-specified token positions. The resulting directions are first-order optimal, explicitly coupling direction extraction with deployment. Experiments reveal that OCGS achieves the highest behavioral-control gains when extraction and intervention operators are matched. Jacobian-lens analysis further shows that distinct steering directions can carry related downstream behavioral content, yet matched OCGS directions yield the strongest layerwise gains in behavior-related readouts. Together, our work shows that the intervention operator should enter the definition of the steering problem itself, with OCGS providing a practical realization of this principle.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.