When Does a Language Model Commit? Answer Preference Before Explicit Answers
Abstract
An answer can be predictable before it is generated even when the model's preference later changes. We measure when preference stabilizes using a language model's continuation probabilities over a specified finite answer set. We define commitment as the first high-margin preference for the eventual answer after which the preferred answer remains unchanged until a specified text event. Across 352 controlled arithmetic responses, preference stabilizes before an explicit judgment. Among 155 responses whose initial preference differs from the final answer, 136 stabilize at the token that completes the printed total or second comparison subtotal. In parity, the outcome word follows four or five tokens later. Initial agreement also differs from stability: 108 of 197 initially agreeing responses later reverse preference. Hidden-state readouts, matched interventions, and held-out stopping comparisons test how this preference is represented and used. Aligning preference with numerical output separates changes in the preferred answer from the time taken to express a judgment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.