acceptodds
Under review as a conference paper at ICLR 2027

Value Safety under Hidden Controllers: Behavior, Specification Exposure, and Residual Interventions

Abstract

Natural-language controllers allow assistants to follow application-specific value priorities, but the instructions can also surface in generated responses. We study value safety as the joint requirement of preserving value-directed behavior while limiting exposure of the internal specification. We introduce a training-free interventions for Value Residual Guidance (VRG), which uses the logit difference between controller-conditioned and base passes through model. VRG-Exact zeroes selected controller-associated residual coordinates and VRG-Soft shrinks them; both retain the base logits rather than banning tokens. Experiments show that stronger guidance can increase literal exposure even when additional behavioral gains are small. Controller text also appears in non-extraction tasks, and intervention effects vary across attack families, models, and steering directions. Additional evaluations show small incremental changes in content safety and task capability for VRG-Soft relative to VRG. Our findings support a partial separation between intended behavioral influence and literal specification exposure, while semantic and functional signals of the configured value remain observable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.