Trajectory-to-Go: Forecasting Changes in Language Model Predictions
Abstract
A language model's intermediate prediction can change little from one step to the next yet still differ substantially from its final output. We ask whether the last two steps help predict how much remains to change. Trajectory-to-Go (T2G) summarizes their magnitudes and changes in direction in both the model's predictions and its hidden states, then learns to forecast disagreement with a chosen later output. Its main readout, Risk-T2G, combines fourteen trajectory and confidence inputs in a two-head error forecast. On held-out Huginn-0125 and two Ouro scales, it reduces error-ranking loss (EAURC) by 7.7–22.3% over confidence-based predictors with matched training objectives and capacity, and outperforms a learned current-state probe in all three models. The fitted forecasts generalize to C4 and LAMBADA without adaptation. Across four ordinary transformers up to 7B parameters, T2G reduces EAURC by 10.8–21.1% over matched confidence. A state-augmented extension improves on either trajectory or state features alone in all eight model–lens settings. Stopping studies evaluate KL-constrained allocation; an answer-change extension tests decision preservation and execution cost. Recent trajectory history thus helps distinguish predictions that look stable from those close to a later output.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.