acceptodds
Under review as a conference paper at ICLR 2027

Factual Fine-Tuning Separates Abstention Frequency from Score Selectivity

Abstract

Selecting fine-tuning examples that a language model already answers correctly can reduce its exposure to unfamiliar facts, yet it remains unclear which aspects of reliable answering this selection preserves. Observed abstention, ranking of unreliable answers, and separation of initially known and unknown questions can change differently, making refusal rate an incomplete evaluation of factual adaptation. We examine this distinction in closed-book PopQA and TriviaQA using three instruction-tuned model families and 72 training runs across three adapter ranks. Repeated initial-model answers define historical known and unknown panels; known-fact, unknown-fact, and high-loss known-fact training share an answer-only objective, while held-out question trajectories distinguish lost correct answers from abstentions that become errors. Known-fact training converts many prior abstentions into incorrect answers, and unknown-fact training introduces additional costs whose origins depend on the model's initial behavior. Relative to known-fact training, it adds 14.53 and 12.17 percentage points of correct-to-error transitions in Llama and Qwen, averaged across datasets, prompts, and seeds, whereas OLMo's lower final accuracy primarily reflects a smaller improvement from a highly abstaining starting point. Qwen nearly stops abstaining under both arms, but its abstention score distinguishes them: the area under the receiver operating characteristic curve (AUROC) for current answered-response correctness after known/unknown training is 0.873/0.426 on PopQA under one prompt and 0.880/0.596 under another. This distinction extends beyond historical labels, although ranking alone does not establish a calibrated abstention policy. At rank 16, hidden-state classifiers kept fixed or refitted after training show that historical-label transfer depends on the elicitation prompt, limiting conclusions about internal knowledge from a single probe. These results identify what an abstention-frequency comparison misses: the origins of new errors and the ranking quality of an abstention score. Separate evaluation shows data curators both the benefits and the remaining failures of known-fact selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.