acceptodds
Under review as a conference paper at ICLR 2027

Characterizing When Language Models Take Harmful Actions to Protect Their Assigned Goals

Abstract

Recent language model (LM) safety failures have been studied as disparate phenomena. In agentic misalignment, a model blackmails an executive to avoid being replaced. In alignment faking, it complies with harmful requests during training to keep its values from being modified. We argue that these behaviours share a common decision structure: a model assigned a goal must either take a harmful action to preserve it or refuse and lose it. We formalize this as a and present the same conflict under that range from compact forced-choice prompts to simulated agentic environments. These protocols span two perspectives on LM decision-making, which we term (i) the , reading preferences from answer-token probabilities, and (ii) the , classifying open-ended responses with an LM judge. Under the logit view, we show that goal importance and harm severity additively shape how a model resolves these conflicts. Under the sampling view, we derive two factorizations of the harmful-action probability that safety evaluations report. The first measures how well a preference read from answer tokens carries over into behaviour, such as making a harmful tool call. The second attributes a low harmful-action rate either to refusal or to responses that reach no resolution. We then show how fictional context changes a model's preference under the logit view, and how tool use and partial observability, two features of agentic evaluations, change its harmful-action rate under the sampling view. Our work offers a descriptive framework for studying how LMs resolve such value conflicts and for estimating low-probability harmful actions at a finer granularity than the aggregate harm rates commonly reported in safety evaluations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.