acceptodds
Under review as a conference paper at ICLR 2027

FamiFT: Familiarity-Adaptive Supervised Fine-Tuning

Abstract

Supervised fine-tuning (SFT) typically applies the same training objective across examples, while recent probability-based methods such as Dynamic Fine-Tuning (DFT) adapt token supervision using the model's online confidence. However, the appropriate strength of such token weighting may itself depend on how well the pretrained model already fits each training example. We first conduct a controlled study and find a consistent familiarity-dependent pattern: relatively familiar examples benefit from sharper probability-based weighting, whereas relatively unfamiliar examples benefit from flatter weighting. Motivated by this observation, we propose Familiarity-Adaptive Fine-Tuning (FamiFT), a simple extension of probability-weighted SFT that assigns each example a sample-specific exponent. FamiFT measures response familiarity once using the frozen initial model, centers it by the dataset median, and uses the resulting score to modulate online token probabilities throughout fine-tuning. Across mathematical reasoning and code generation, FamiFT consistently improves average performance over standard SFT and competitive token-weighting baselines, providing a simple way to incorporate pretrained-model familiarity into supervised fine-tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.