A Simple Statistic for Discovery Capacity in Language Models
Abstract
Large language models (LLMs) show growing promise for scientific discovery, yet basic questions remain about what enables reliable discovery. Can a simple statistic of the training data reveal a model's capacity for discovery? We study this question through missing mass and develop a theoretical account in which monofacts–the facts observed once during training–set a budget for discovery. In materials science and regulatory genomics, we directly manipulate monofact prevalence and find that more monofacts consistently yield more discoveries across two open-weight model families, with and without reasoning traces, and across sampling strategies and reasoning budgets. Further experiments isolate the effect of monofact prevalence from key confounds, show that higher monofact prevalence increases mass placed on unobserved outputs, and find that stronger conditioning can substantially increase discovery. Together, our results identify monofact prevalence in training data as a key factor governing LLM discovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.