acceptodds
Under review as a conference paper at ICLR 2027

Sharp at the Peak, Rich in the Prior: Peak-Preserving Prior Distillation for RL-Ready Large Language Models

Abstract

Supervised fine-tuning (SFT) for large language models should acquire demonstrated behavior while retaining useful pretrained preferences and alternatives for subsequent reinforcement learning (RL). Conventional one-hot supervision promotes the demonstrated token without specifying which relationships among alternative continuations should be preserved. We introduce **Peak-Preserving Prior Distillation ()**, a principled approach that specifies the desired supervisory distribution before optimization. Guided by **Greedy Guarantee** and **Minimal Intervention**, constructs a closed-form target that promotes demonstrated-token dominance while preserving relative probabilities among retained non-target alternatives from the current student. This directly embeds acquisition and retention into the supervision target, providing dense feedback without teacher probability vectors or an additional teacher forward pass. Experiments across three models from two families on mathematical and medical reasoning demonstrate stronger SFT and subsequent GRPO performance alongside strong general capabilities. Compared with the respective HardSFT baselines, domain-averaged greedy accuracy improves by up to 6.1 percentage points after SFT and 10.3 points after GRPO. Subsequent analysis further support effective task adaptation with limited predictive drift and retained prior preferences. These findings highlight principled target design as a foundation for learning demonstrated behavior while preserving the capacity for further reinforcement learning. Our code has been open-sourced at the anonymous link https://anonymous.4open.science/P3D-433F.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.