DRAFT: Disentangling Retrieval and Abstention via Factual Post-Training
Abstract
Factual hallucination in large language models (LLMs) manifests along two axes: the model may fail to a fact it has been taught, or it may an answer for a fact it was never taught. Reducing hallucination therefore requires two complementary skills: accurate retrieval and appropriate abstention. The standard approach is post-training, where supervised fine-tuning (SFT) embeds knowledge and reinforcement learning (RL) aligns behavior, yet how these stages interact for factuality is not well understood. We present a systematic study under a controlled setup that probes a model's knowledge boundary and isolates SFT and RL effects on retrieval and abstention within one pipeline. We find: RL with binary rewards reliably teaches abstention by enforcing a designated answer/abstain boundary, and the form of SFT data determines what RL can subsequently improve. Together, these findings reveal that RL acts as elicitation over SFT-encoded representations rather than acquiring new knowledge. SFT data construction determines what the model can know, while RL determines how reliably it retrieves or abstains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.