acceptodds
Under review as a conference paper at ICLR 2027

Steer Responsibly: Towards A Grounded Library of Interpretable Steering Vectors

Abstract

Can concepts grounded in human feedback become reliable activation-steering controls? We build Steer Responsibly to test this path: discover signed preference features, independently confirm their descriptions, extract model-specific vectors, and qualify their behavioral effects against quality and control criteria. Twelve of 32 features pass automated semantic confirmation, yielding 48 candidate tensors across Qwen3-8B and Gemma-3-4B-IT, each tested in both intervention directions. None meets the behavioral qualification criteria. Of 96 signed candidates, 86 select the unsteered fallback. Among ten held-out interventions, five miss the target margin in their means, two more fail random-control specificity in their means, and three meet all mean margins but fail uncertainty bounds. Concrete case studies expose limits on attribution: 36 signed candidates have no baseline score headroom; a random direction better expresses a supportive cue than an empathy vector; and near-identical answers receive opposite labels. A separate public-vector audit also finds order-sensitive judgments. We diagnose the gap between semantic grounding and qualified control. The findings make measurement validity part of the library itself to support future research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.