acceptodds
Under review as a conference paper at ICLR 2027

Equal Gain per Unit Statistical Distance: A Principle for Token Reweighting

Abstract

Supervised fine-tuning (SFT) offers a simple way to learn from fixed demonstrations, but improving its generalization remains a challenge. Recent methods assign token-level weights to the cross-entropy loss and report empirical gains, yet benchmark results alone do not explain how these weights value learning progress. In this work, we compare token weights in a common unit: the objective gain each weight assigns per unit of Fisher–Rao (FR) distance toward the demonstrated token. We find that Dynamic Fine-Tuning (DFT) and InfoSFT each assign more uniform gains than SFT over different ranges of target-token probability, offering a possible explanation for their empirical advantages. Building on this insight, we derive Vocab Relative Advantage (VRA), the token weight that gives uniform gain per unit of FR distance across the probability range and is unique up to scale. VRA achieves higher average accuracy than both SFT and DFT on each of six backbones. On Qwen3-8B, it reaches 39.23% across five mathematics benchmarks, compared with 29.30% for SFT, and also surpasses DFT, InfoSFT and ProFit. Matched one-step comparisons further show that, relative to SFT, VRA redistributes progress from low-confidence targets to middle- and high-confidence targets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.