acceptodds
Under review as a conference paper at ICLR 2027

Confidence-Gated Step Selection For Long-Cot Supervised Fine-Tuning

Abstract

Supervised fine-tuning (SFT) on long chain-of-thought (CoT) reasoning traces is an effective way to improve reasoning models. Standard SFT applies one-hot supervision to every sampled token, even in steps where current model confidence exceeds the reference score from the trace generator. Using distribution matching as an analytical reference, we show that hard- and soft-target supervision can yield opposite gradient signs for the sampled token’s logit. This observation motivates selecting supervision according to the model’s evolving confidence. We propose Confidence-Gated Step Selection (CGSS), a selective SFT method that dynamically masks reasoning-step losses using the gap between reference scores and current model confidence. Reference scores are precomputed for each step, while current scores reuse the training forward pass. The criterion operates on a fixed corpus of reasoning traces and requires neither token-level alignment nor full- vocabulary teacher distributions during SFT. Experiments across five benchmarks show improvements over vanilla SFT and heuristic selection baselines, achieving an average accuracy gain of up to 6.3 percentage points while reducing step-level confidence discrepancies, with no additional inference-time computatio

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.