SVM-STEER: Steer Clear of Biases with a Little Support
Abstract
Activation steering enables targeted control of large language model (LLM) behavior through interventions in internal representations. Existing approaches primarily construct steering vectors using unsupervised methods and often rely on brute-force searches to identify intervention layers. In this work, we propose a supervised alternative based on Support Vector Machines (SVMs). We evaluate our approach across two case studies spanning bias mitigation and censorship control, using six datasets and five open-weight LLMs. Our method reduces gender bias by an average of 29.3% and racial bias by 39.7% across models and steering options. For censorship evasion, steering decreases refusal probabilities by an average of 65% . For censorship enhancement, it increases the average refusal probability on XSTest from 0.064 to 0.83. We further show that steering effectiveness varies across model families and identify mode collapse under large steering coefficients, highlighting a trade-off between intervention strength and behavioral control. Our results demonstrate that supervised steering provides a simple, effective, and broadly applicable framework for controlling model behavior through internal representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.