WordNet-Guided Domain Data Selection for LLM Continual Pre-training
Abstract
Continual pre-training (CPT) adapts language models to specialized domains, but selecting useful training data from large domain corpora remains difficult. We introduce Frequency-Conceptual Difficulty (FCD), a model-free metric that combines lexical frequency with semantic distance in WordNet to score document difficulty. In the medical domain, FCD-Hard selects a fixed one-third subset and trains on it for three epochs. Under the same total token exposure, this strategy improves domain-average accuracy over full-corpus CPT by 1.36 percentage points and outperforms the compared data-selection baselines. We also examine whether the preferred FCD range is domain dependent. Medical performance is highest when training on documents with rare, WordNet-distant terminology, whereas FCD-Easy performs best in our experiment on an English translation of FinEval. Qualitative examples suggest that some high-FCD documents in the financial corpus contain technical vocabulary from other domains. These results support using FCD as a lexical criterion for domain-data analysis while showing that the sampling direction must be selected for the target corpus.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.