acceptodds
Under review as a conference paper at ICLR 2027

Same Candidates, Stronger Reasoning: Rethinking What Post-Training Learns

Abstract

Post-training transforms pretrained language models into notably stronger reasoners. Yet recent evidence suggests that the changes driving these gains are concentrated at a small number of token-level decisions. This raises a fundamental puzzle: at these consequential decisions, does post-training rely on candidates beyond the pretrained model's high-ranked set, or primarily reorganize preferences among them? Across 31 model–task–checkpoint settings, 92.9% of post-training Top-1 changes remain within the pretrained Top-16 on average, while their relative preferences are substantially reorganized. More strikingly, this inheritance persists even on successful post-training-induced prefixes that are not visited by the pretrained policy. Accumulated local preference changes can therefore compose into genuinely new global competence. Motivated by this finding, we introduce ProQ, a post-training paradigm that shifts the learning target from full-distribution next-token adaptation to value-guided selection over the model's own proposals. Within a single LLM and shared vocabulary, standard-token logits retain normal reasoning and candidate proposal, regularized toward the pretrained model; reserved-token logits are repurposed to produce candidate-conditioned value scores. These scores provide a natural RL interface for local selection. Trained only on mathematical reasoning, ProQ achieves strong gains on math benchmarks and transfers to out-of-domain coding and knowledge tasks; the benefits persist across different model sizes and families.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.