acceptodds
Under review as a conference paper at ICLR 2027

When Does a Language Model Commit? Answer Preference Before Explicit Answers

Abstract

An answer can be predictable before it is generated even when the model's preference later changes. We measure when preference stabilizes using a language model's continuation probabilities over a specified finite answer set. We define commitment as the first high-margin preference for the eventual answer after which the preferred answer remains unchanged until a specified text event. Across 352 controlled arithmetic responses, preference stabilizes before an explicit judgment. Among 155 responses whose initial preference differs from the final answer, 136 stabilize at the token that completes the printed total or second comparison subtotal. In parity, the outcome word follows four or five tokens later. Initial agreement also differs from stability: 108 of 197 initially agreeing responses later reverse preference. Hidden-state readouts, matched interventions, and held-out stopping comparisons test how this preference is represented and used. Aligning preference with numerical output separates changes in the preferred answer from the time taken to express a judgment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.