acceptodds
Under review as a conference paper at ICLR 2027

BREW: Progress-Aware Token Weighting for Open-Ended Writing

Abstract

Supervised fine-tuning (SFT) for open-ended writing learns from a single recorded realization even where several continuations are plausible. Under uniform SFT, every supervised token keeps weight one throughout training, so continued fitting concentrates probability on the recorded realization even after the token becomes predictable, a phenomenon we call single-reference overcommitment. Existing confidence-based weighting partly eases this problem by downweighting low-confidence tokens, but its weight can keep rising as the target becomes more predictable, leaving no signal to reduce weighting after substantial progress. We introduce BREW (Base-Relative Weighting), which compares each token's current loss with its loss under the frozen base model. The current loss measures how learnable the token is now, while the base-relative ratio measures progress from its own starting point. This keeps the cautious entry of confidence weighting while adding a progress-aware exit. In the primary setting, BREW performs best on every writing metric, improving WritingBench by 2.1 and LongBench-Write by 3.7 over the strongest baseline. Its WritingBench advantage holds across models and datasets, and the same configuration remains competitive on answer-determined mathematic tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.