Scaling LLM Knowledge Boundaries via Distribution-Optimized Synthesis
Abstract
Knowledge injection via synthetic data is crucial for enhancing Large Language Models (LLMs). However, current synthesis methods simply stop at preset token counts or fixed data ratios, lacking awareness of knowledge distribution. This results in some knowledge points being sparse while others are redundant, limiting LLM knowledge boundaries. We revisit knowledge injection from a distribution perspective and hypothesize that an optimal knowledge density range exists for maximizing knowledge boundary expansion. We propose **KDoS** (**K**nowledge **D**istribution-**o**ptimized **S**ynthesis), a framework that introduces knowledge density as a controllable variable to quantify the concentration–dispersion degree of knowledge in semantic space, and uses a three-stage feedback mechanism to drive synthesis, shifting from blind generation to distribution-optimized synthesis. We construct Wikipedia-based synthetic data with varying knowledge distributions and conduct experiments on models from 0.6B to 16B (Qwen, Ling, LLaMA) and data scales from 1B to 5B tokens. Our key findings are: (1) an optimal knowledge density range consistently maximizes boundary expansion; (2) this range is stable across backbones and scales; (3) KDoS outperforms representative synthesis baselines across six factual knowledge benchmarks. Our work offers a new perspective and practical framework for synthetic data-driven knowledge injection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.