Value Safety under Hidden Controllers: Behavior, Specification Exposure, and Residual Interventions
Abstract
Natural-language controllers allow assistants to follow application-specific value priorities, but the instructions can also surface in generated responses. We study value safety as the joint requirement of preserving value-directed behavior while limiting exposure of the internal specification. We introduce a training-free interventions for Value Residual Guidance (VRG), which uses the logit difference between controller-conditioned and base passes through model. VRG-Exact zeroes selected controller-associated residual coordinates and VRG-Soft shrinks them; both retain the base logits rather than banning tokens. Experiments show that stronger guidance can increase literal exposure even when additional behavioral gains are small. Controller text also appears in non-extraction tasks, and intervention effects vary across attack families, models, and steering directions. Additional evaluations show small incremental changes in content safety and task capability for VRG-Soft relative to VRG. Our findings support a partial separation between intended behavioral influence and literal specification exposure, while semantic and functional signals of the configured value remain observable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.