acceptodds
Under review as a conference paper at ICLR 2027

Not All Conflicts Are Equal: Forecasting Token Learning Trajectories for Supervised Fine-Tuning

Abstract

Supervised fine-tuning adapts pretrained language models to new domains, but often degrade their general capabilities. We study conflicting tokens, where the supervised target disagrees with a confident pretrained prediction. Recent token-level methods treat these tokens differently depending on the goal: suppressing them better preserves general capabilities but often believed to limit adaptation, while emphasizing them improves target-domain adaptation at the cost of retention. However, achieving strong target-domain adaptation while preserving general capabilities remains challenging. This motivates us to look more closely at conflicting tokens. We find that they do not behave uniformly during SFT: some resolve quickly but often revert, some are learned gradually, and others remain difficult to learn. Retaining gradually learned conflicts while masking other two groups gives the best adaptation-retention balance among the strategies we study. Based on this finding, we introduce ForeToken, an efficient pre-SFT forecasting method that estimates each conflict's resolution tendency from the initial optimization direction. This avoids an extra SFT run to observe token trajectories before constructing the mask. Across mathematical, medical, agentic, and code domains and three model families, ForeToken improves target-domain performance while better preserving general capabilities. On Qwen3-8B, it outperforms the strongest token-level baseline by 7.6 points on mathematical reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.