Innovation Prior: Learning to Discover by Distilling the Scientist's Mind
Abstract
Large language models (LLMs) excel at tasks that ask for known methods to be applied well, but on research tasks they recombine and retune existing methods instead of inventing new ones as human scientists do. A harness around a frozen model can only select among what the model produces, and reinforcement learning can only strengthen what it already does, so the missing ability has to come from training data. Papers are in that data, but a paper presents a finished result and never shows how its authors thought. We design a pipeline that reconstructs that thinking and use it to build *Innovation Prior*, a dataset of how scientists reasoned their way to their methods. Each reconstruction draws on multiple sources beyond the paper and includes the authors' exploration of routes that never reached the final method. We fine-tune Qwen3.5-4B and 9B on *Innovation Prior*, then run one RL recipe from the fine-tuned model and from the base model. RL from the fine-tuned model outperforms RL from the base on research benchmarks at both sizes, while none of our training data targets these benchmarks. *Innovation Prior* thus sets a prior for discovery that RL then amplifies. Ablations show that both the multiple sources and this exploration are necessary. The advantage carries over to test-time scaling, where the model scores higher for the same compute, and generalizes to research judgment. The model also explores more broadly, both within and across attempts, and case studies show that it reaches simple but effective methods from first principles.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.