acceptodds
Under review as a conference paper at ICLR 2027

Goal Encoding in LLMs

Abstract

We investigate whether and how LLMs encode goals: internal neural representations that precede and guide action, that are flexible enough to be realized through many different execution paths. We construct a coding setting with multi-step tasks that allows us to evaluate goal-seeking behavior and the selection of sub-goals. In this setting, we localize goal-seeking causal effects in the output of a sparse set of attention heads. The same heads can be used to decode the model's future actions in natural language: e.g., "I am counting unique values,” or "I'm checking if a word is palindrome.” We verify that this encoding is causal for model's next action, persists across long token sequences, and is plastic: transplanting it into a different environment (e.g., another programming language) triggers the same operation, realized with a different sequence of tokens. Although these heads are extracted using a coding task, we find that their goal-seeking and goal-verbalization properties generalize to natural text and few-shot ICL settings. Finally, we discuss how transformer LMs, whose computation is bounded by layer depth, can maintain such a goal encoding across multiple tokens through frequent re-computations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.