acceptodds
Under review as a conference paper at ICLR 2027

When Fine-Tuning Gains Reverse: Separating Prefix and Continuation Effects

Abstract

Supervised fine-tuning (SFT) can improve accuracy at a short output budget yet reduce it at a longer one, relative to the same initial checkpoint. To locate the change, we cross initial and fine-tuned prefix sources with both continuation models, comparing one stage while holding the other fixed. Our models are adapted with LoRA on 500 GSM8K training examples. On 50 generated two-variable linear systems, three Llama runs outperform the initial checkpoint at 192 tokens but underperform it at 768; the larger loss follows the prefix source. On all 1,319 GSM8K test questions, the tested Llama run again shows a prefix-driven loss, whereas three Qwen runs improve mainly through continuation. We then remove terminal end-of-sequence (EOS) targets during training. Exchanging the learned stopping probabilities and conditional content distributions assigns 74.7–81.1% of the algebra ablation gain to stopping. EOS ablation also improves MATH-500 and HumanEval accuracy, although recovery is incomplete in some settings. Checkpoint comparisons should therefore match the intended output budget; crossed continuations help identify which stage changed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.