acceptodds
Under review as a conference paper at ICLR 2027

What Happens to the Goal? Distinct Failures of Policy Alignment in Goal Misgeneralization

Abstract

Goal misgeneralization is a common failure mode in which an agent behaves competently out of distribution but does not pursue the intended goal. From the outside these failures look alike, yet we still understand little about what happens to the intended goal inside misgeneralized agents. We trace the goal through agents in three reinforcement-learning environments, asking whether it is still represented, whether the policy reads it out, and what organizes behavior instead. We find three distinct failures of policy alignment, in which the goal gets lost at different depths: (1) functional target substitution, where the goal is inert to the policy and behavior follows an alternative target, (2) a policy shortcut, where the goal is represented but not used as a goal and behavior follows a learned action rule, and (3) cue competition, where the goal reaches the decision but is outweighed by a proxy cue. Where the goal gets lost explains which recovery works. Under cue competition, steering along the axis on which the two cues compete raises intended-goal reach from to , while the shortcut recovers only after learning a new readout, from to , and under target substitution neither restores goal-directed behavior. So goal misgeneralization is not one failure, and recovering the goal requires knowing where it got lost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.