When Local Choices Matter in Language Model Generation
Abstract
Language models can reach the same answer through different plans, and produce different answers from the same plan. To explain their errors and decide where to intervene, we need to know which choices influence the outcome. Differences between sampled answers mix the effect of a choice with randomness in later generation. We introduce outcome-resolution covariance (ORC), a matrix describing how future outcome probabilities vary across choices under a fixed continuation policy. Independent continuations separate choice effects from continuation noise and yield unbiased ORC estimates. We use one continuation to define intervention weights and another to predict their gain, separating effects of the shared choice from dependence within a continuation. Across three Qwen text-to-SQL panels, these forecasts more accurately predict gains on fresh continuations, including new questions and plans. On a 128-question panel, they reduce mean squared error by 80% relative to uncalibrated same-continuation forecasts; using them to allocate a fixed intervention budget offline improves correctness by 1.05 percentage points over equal-budget random selection. Recording outcome categories at different checkpoints reveals plan-specific failures that correctness scores hide and state differences that later computation removes. A covariance identity connects reweighting gains to probability shifts, branch variation, and target alignment. Ledger experiments show why alignment matters: on tasks with incorrect consensus targets, reweighting increases agreement while reducing correctness on average. Variation between branches provides the room for steering; alignment with the task determines the benefit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.