acceptodds
Under review as a conference paper at ICLR 2027

SVM-STEER: Steer Clear of Biases with a Little Support

Abstract

Activation steering enables targeted control of large language model (LLM) behavior through interventions in internal representations. Existing approaches primarily construct steering vectors using unsupervised methods and often rely on brute-force searches to identify intervention layers. In this work, we propose a supervised alternative based on Support Vector Machines (SVMs). We evaluate our approach across two case studies spanning bias mitigation and censorship control, using six datasets and five open-weight LLMs. Our method reduces gender bias by an average of 29.3% and racial bias by 39.7% across models and steering options. For censorship evasion, steering decreases refusal probabilities by an average of 65% . For censorship enhancement, it increases the average refusal probability on XSTest from 0.064 to 0.83. We further show that steering effectiveness varies across model families and identify mode collapse under large steering coefficients, highlighting a trade-off between intervention strength and behavioral control. Our results demonstrate that supervised steering provides a simple, effective, and broadly applicable framework for controlling model behavior through internal representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.