Data-Efficient Narrative Knowledge Injection via Entity-Level Diffusion
Abstract
Large language models often struggle to acquire knowledge directly from narrative corpora, while existing approaches rely on extensive synthetic augmentation to make the injected knowledge accessible. We attribute this difficulty to a binding–access mismatch: autoregressive training binds knowledge to highly specific, left-to-right contexts, whereas inference-time queries may provide substantially different cues. We introduce EntDiff, a data-efficient approach that learns primarily from source narratives with only lightweight preprocessing. Compared with augmentation-based methods, EntDiff generates more than three orders of magnitude fewer tokens during data preparation and trains on 2.5–6 fewer tokens. EntDiff reshapes how narrative knowledge is bound during training by making entity references explicit, limiting dependence on overly specific contexts, and creating diverse bidirectional associations accessible from different inference-time cues. Experiments show that EntDiff matches or outperforms strong augmentation-based baselines on dataset-paired question answering while substantially improving episodic recall.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.