Teaching LLMs New Knowledge Without Hallucinations
Abstract
Large language models increasingly operate in settings where they encounter information that was unavailable during pretraining. While such knowledge can be stored externally, continually internalizing it into model weights could improve efficiency and enable reasoning across information acquired at different times. We study knowledge acquisition from natural-language documents and identify a central failure mode: fine-tuning on new knowledge can rapidly increase hallucinations. Models become more willing to answer questions about the newly learned domain even when the required information was never provided, and newly acquired facts can also spill over into answers to unrelated questions. Such hallucinations are undesirable in themselves and are especially concerning in continual-learning systems, where they may accumulate over time. We introduce hallucination-aware training, which combines context distillation with hallucination-probing questions that teach the boundary of the acquired knowledge and a regularization objective that preserves the model’s behavior on unrelated questions. Across Qwen3-32B, GPT-OSS-20B, Gemma4-31B, and Qwen3.8-27B, our approach substantially reduces hallucinations and knowledge spillover while preserving strong acquisition of new knowledge. We further find that models can compose facts learned independently from different documents, even though training processes only one document at a time and provides no explicit supervision for cross-document composition. Our results suggest that LLMs can internalize new knowledge while learning both how to use that knowledge and where its boundaries lie.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.