Longer Answers, Heavier Questions: How Loss Weighting Shapes Comparisons of Synthetic Reasoning Data
Abstract
When a language model is fine-tuned on synthetic reasoning data with a token-mean loss, each question counts in proportion to the length of its target. We call this implicit question weighting. A recipe that sets how much reasoning each question receives therefore also sets how much each question counts in training. With targets of similar length the weights are close to uniform; in the mixed-depth corpus we study, direct-answer targets are 42% of the questions but receive under 6% of the loss weight. We measure how strongly a comparison between two recipes depends on this weighting. Both recipes come from one teacher pool on the same 11,853 questions: one gives every question a brief explanation, the other mixes direct answers with deeper reasoning. With the targets held fixed, replacing the token mean by an equal weight per example causes no detectable change for the fixed-depth corpus but costs the mixed-depth corpus 8.6 and 5.5 points of four-benchmark accuracy on Qwen3-4B-Base and OLMo-3-7B (three seeds per arm), so the lead of the fixed-depth corpus grows from 1.6 to 9.9 points on one student and from 3.1 to 8.3 on the other. Training the fixed-depth targets with the mixture’s own question weights does not reproduce this response, so the weights alone do not explain it. With each source dataset’s share of the loss held fixed, moving the direct-answer targets to their share of the examples costs more than equalizing the examples within the direct-answer and the reasoning groups. Comparisons of reasoning datasets should therefore report their loss objective and the loss share of their shortest targets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.