Steer Responsibly: Towards A Grounded Library of Interpretable Steering Vectors
Abstract
Can concepts grounded in human feedback become reliable activation-steering controls? We build Steer Responsibly to test this path: discover signed preference features, independently confirm their descriptions, extract model-specific vectors, and qualify their behavioral effects against quality and control criteria. Twelve of 32 features pass automated semantic confirmation, yielding 48 candidate tensors across Qwen3-8B and Gemma-3-4B-IT, each tested in both intervention directions. None meets the behavioral qualification criteria. Of 96 signed candidates, 86 select the unsteered fallback. Among ten held-out interventions, five miss the target margin in their means, two more fail random-control specificity in their means, and three meet all mean margins but fail uncertainty bounds. Concrete case studies expose limits on attribution: 36 signed candidates have no baseline score headroom; a random direction better expresses a supportive cue than an empathy vector; and near-identical answers receive opposite labels. A separate public-vector audit also finds order-sensitive judgments. We diagnose the gap between semantic grounding and qualified control. The findings make measurement validity part of the library itself to support future research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.