acceptodds
Under review as a conference paper at ICLR 2027

From Principles to Priorities: Behavior-Relative Latent Adaptation under Goal Conflict

Abstract

As language models act autonomously over longer horizons, alignment depends on whether a principle governs action when it conflicts with task completion. A model may know a harm-avoidance rule yet fail to prioritize it under goal conflict, or apply it so broadly that legitimate actions are suppressed. We study whether a compact set of 37 constitution-derived contrastive stems can induce a persistent harm-avoidance priority that transfers to new scenarios. We introduce BEHAVE, a behavior-relative latent adaptation procedure that uses a model-state-dependent preference-flip threshold as a reference for training pressure and refreshes this reference during adaptation. Across three models, the taught rule transfers from first-person completions to new operational decisions. The shift largely preserves performance on conventional capability benchmarks, but benign-action retention depends strongly on the base model. In a multi-turn tool-use evaluation, BEHAVE reduces selection of a harmful tool under both explicitly harmful and benign-sounding names, while sensitivity to the tool name remains. These results show that compact principle-derived supervision can induce a decision priority that persists beyond the adaptation examples and transfers across substantially different decision settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.