Today's Targets, Tomorrow's Context: Measuring Token-Level Future Influence in SFT
Abstract
Reasoning supervised fine-tuning (SFT) typically filters entire trajectories based on final-answer correctness, yet applies direct supervision to every response token in each retained trajectory. Consequently, trajectory-level verification cannot determine whether every token is a suitable training target. In an autoregressive model, the same token serves first as a prediction target and, once generated, as context for subsequent predictions. Directly supervising a token increases the model's tendency to generate it in the future. However, even in a trajectory that reaches the correct answer, some tokens may, as context, make the fixed demonstrated continuation harder for the current model to predict. Motivated by this observation, we introduce Future Influence, which measures the signed effect of each response token as context on subsequent prediction and uses this direction to select token-level supervision. Specifically, we attenuate the information a token transmits to later positions and determine the direction of its effect from the resulting change in prediction loss on the fixed demonstrated continuation. Experiments show that negatively directed tokens are common even in correct reasoning trajectories. Supervision selected by Future Influence consistently improves reasoning performance across model scales, training datasets, and evaluations. Our findings show that final-answer correctness does not imply that every token in a trajectory should be directly reinforced, and that accounting for tokens' roles as future context provides a new perspective on constructing more effective supervision for SFT.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.