Locks and Keys: Learning Scientific Inspiration Retrieval as Pattern Matching at Scale
Abstract
A new hypothesis often arises when a research problem meets complementary knowledge not previously associated with it. Making such discoveries thus begins with inspiration retrieval: finding, among the vast body of existing knowledge, the piece that fits the problem. The task looks unlearnable, since discoveries are new connections by definition and the ones a model must find at test time are absent from its training data. Yet it is learnable, and we propose why: inspiration retrieval is essentially pattern matching. A research problem is a lock and an inspiration is a key; individual pairs are new, but the way a problem's structure fits an inspiration's structure recurs, so the rule can be learned from records of past discoveries by ordinary supervised fine-tuning. Experiments bear this out: what a trained model learns is the fit, not a memory of similar pairs and not resemblance between problem and paper. The view also says how to learn the task: as a score on each problem–paper pair, computed by embedding and reranking models, rather than as a selection over a candidate set by a generative reasoning model, as in prior large-scale training. Given the same base model and annotations, scoring pairs is far more accurate than generative selection, at a fraction of the cost, and stays ahead as accuracy grows log-linearly over hundreds of thousands of training pairs. Finally, the learned fit uses only a small fraction of the embedding dimensions available, and the fraction does not increase with more training data: the task is limited not by representation capacity but by how much of the fit has been learned.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.