acceptodds
Under review as a conference paper at ICLR 2027

Learnable Reference Fine-Tuning for Rollout-Free Mathematical Reasoning

Abstract

Reinforcement learning can improve mathematical reasoning in large language models, but large-scale online rollouts are expensive. We study rollout-free fine-tuning from verified solutions alone, where the absence of negative examples leaves the contribution of each demonstration difficult to calibrate. We propose Learnable Reference Fine-Tuning(LRFT), which jointly trains a target policy and a reference policy on verified demonstrations. The reference models the demonstration distribution, and target-to-reference likelihood ratios modulate token-level target updates. Unlike target-only self-weighting or fixed-reference regularization, LRFT learns its update scale from the verified data. Across four checkpoints from two Qwen families and five mathematical-reasoning benchmarks, LRFT achieves the highest mean Pass@1 among all evaluated methods for every checkpoint and the best performance in most reported Pass@k comparisons. Relative to DFT, it also shows lower measured KL divergence from the base policy and less polarized gold-token probabilities. These results suggest that a demonstration-fitted reference provides an effective rollout-free learning signal while limiting distributional shift.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.