acceptodds
Under review as a conference paper at ICLR 2027

DRAFT: Disentangling Retrieval and Abstention via Factual Post-Training

Abstract

Factual hallucination in large language models (LLMs) manifests along two axes: the model may fail to a fact it has been taught, or it may an answer for a fact it was never taught. Reducing hallucination therefore requires two complementary skills: accurate retrieval and appropriate abstention. The standard approach is post-training, where supervised fine-tuning (SFT) embeds knowledge and reinforcement learning (RL) aligns behavior, yet how these stages interact for factuality is not well understood. We present a systematic study under a controlled setup that probes a model's knowledge boundary and isolates SFT and RL effects on retrieval and abstention within one pipeline. We find: RL with binary rewards reliably teaches abstention by enforcing a designated answer/abstain boundary, and the form of SFT data determines what RL can subsequently improve. Together, these findings reveal that RL acts as elicitation over SFT-encoded representations rather than acquiring new knowledge. SFT data construction determines what the model can know, while RL determines how reliably it retrieves or abstains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.