acceptodds
Under review as a conference paper at ICLR 2027

Reference-Aware Dynamic Fine-Tuning for Multi-Sample Reasoning

Abstract

Dynamic Fine-Tuning (DFT) has recently shown that a simple detached reweighting of supervised fine-tuning can substantially improve average sampled accuracy on reasoning tasks. We identify a complementary limitation: the same DFT checkpoints that improve average-sample performance can degrade high-k pass rates, sometimes falling below the base model under multi-sample decoding. We trace this tension to a local distributional effect of the DFT update. By increasing the demonstrated-token probability, DFT also contracts high-probability non-reference alternatives through the softmax; this sharpening is beneficial when alternatives are local noise, but can reduce pass@k when alternatives correspond to viable reasoning continuations. Motivated by this mechanism, we propose Reference-Aware Dynamic Fine-Tuning (RA-DFT), an objective that uses the frozen base model's teacher-forced probability as a local confidence signal. RA-DFT attenuates DFT-style sharpening at low-reference-probability positions, where alternatives are more likely to support multi-sample coverage, while retaining strong sharpening at high-reference-probability positions. The method preserves the standard supervised fine-tuning pipeline and only changes the token-level loss weight. Across five models and six mathematical reasoning benchmarks, RA-DFT consistently improves over DFT across the evaluated pass@k range and improves majority voting, while preserving the average-sample gains that make DFT a strong baseline. Compared with a DFT+KL anchoring baseline, RA-DFT yields a better overall low-k/high-k balance, suggesting an advantage of reference-conditioned selective sharpening over global base anchoring.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.