In the Shadow of Reasoning: Exploring Behavioral Steering in Large Reasoning Models
Abstract
In large reasoning models (LRMs), activation steering effectively controls chain-of-thought (CoT) progress but has limited success in behavioral steering, such as refusal. This study reveals that this disparity arises from the activation geometry of LRMs' CoT. During reasoning, the CoT-progress direction captures the dominant variation in CoT activations, while the weaker refusal signal remains in its shadow. This geometry poses challenges for refusal steering in both identifying the refusal direction and controlling intervention strength. To validate these insights, we propose Substeer, which extracts a clear refusal direction and performs dynamic strength control for effective refusal steering while preserving CoT coherence. Experiments on four LRMs and six jailbreak attacks demonstrate the effectiveness of Substeer. Further evaluation of other behavioral steering tasks, such as myopic reward, supports the generalizability of our findings. We open-source our code and data at https://anonymous.4open.science/r/Substeer/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.