From Principles to Priorities: Behavior-Relative Latent Adaptation under Goal Conflict
Abstract
As language models act autonomously over longer horizons, alignment depends on whether a principle governs action when it conflicts with task completion. A model may know a harm-avoidance rule yet fail to prioritize it under goal conflict, or apply it so broadly that legitimate actions are suppressed. We study whether a compact set of 37 constitution-derived contrastive stems can induce a persistent harm-avoidance priority that transfers to new scenarios. We introduce BEHAVE, a behavior-relative latent adaptation procedure that uses a model-state-dependent preference-flip threshold as a reference for training pressure and refreshes this reference during adaptation. Across three models, the taught rule transfers from first-person completions to new operational decisions. The shift largely preserves performance on conventional capability benchmarks, but benign-action retention depends strongly on the base model. In a multi-turn tool-use evaluation, BEHAVE reduces selection of a harmful tool under both explicitly harmful and benign-sounding names, while sensitivity to the tool name remains. These results show that compact principle-derived supervision can induce a decision priority that persists beyond the adaptation examples and transfers across substantially different decision settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.