acceptodds
Under review as a conference paper at ICLR 2027

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

Abstract

A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The _power distribution_ assigns each complete answer a probability proportional to the model's probability of that answer raised to a power greater than one, which moves probability toward the answers the model considers most likely. We call this sharpening. Sampling from the power distribution improves reasoning accuracy without changing the model's parameters, but it requires generating and scoring many candidate answers per query. We show that a model can instead be trained once to produce such answers in one ordinary generation. Our method, on-policy power distillation (OPPD), runs a sequential Monte Carlo sampler in which the model being trained generates candidate answers and a frozen teacher's power distribution weights them. The same teacher probabilities give each complete answer a training weight for weighted maximum likelihood. Training raises single-generation accuracy by up to points on MATH500 and on GSM8K over the untrained model at the same temperature, and one generation of the trained model scores and points above published power sampling with candidates, obtaining 94% of the gain that candidates give the untrained model. To place the size of this gain in context, we also train GRPO, a reinforcement-learning method that rewards answers verified as correct, from the same checkpoint with the same budget: oppd scores , and points higher on MATH500, GSM8K and AIME while using no reference answers. The two methods use different information and are complementary rather than alternatives, and applied after GRPO OPPD adds up to points more. Although it trains only on mathematics questions, OPPD raises HumanEval code accuracy by up to points. We measure the sharpening a model absorbs as an exponent fitted between its token probabilities and those of the untrained model. One loss coefficient moves this exponent between and , against for ordinary on-policy distillation, and it rises mostly on the answers the model writes itself. The gains hold across model families and sizes, including a model already trained with verified rewards, on which lowering the sampling temperature gives nothing while OPPD adds points on MATH500.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.