Training LLMs to Guide Research by Studying Scientific Papers
Abstract
Training language models to make progress on open-ended research problems is difficult because useful feedback is hard to attribute to individual research decisions, such as proposing an idea, developing a method, or designing an experiment. Directly specifying rewards for such intermediate decisions can also become overly prescriptive, narrowing what counts as a promising research direction. In AI research, even obtaining sufficient reliable feedback is quite expensive. We introduce PESTO, or "Paper-Extracted Surrogate Task Optimization", an approach for improving a model's ability to conduct open-ended AI research by deriving focused research tasks from scientific papers. We decompose papers into connected decisions such as identifying literature gaps, proposing methods, designing experiments, and interpreting results, and turn them into shorter-horizon problems with criterion-level rubric rewards. Applying this construction across many papers makes individual decisions easier to evaluate while preserving a broad range of research approaches. We train a 27B-parameter model with RL on tasks derived from ICML, ICLR, NeurIPS, and arXiv papers. Training improves performance on held-out research tasks and transfers to more open-ended settings where proposed ideas need to be implemented and evaluated. PESTO improves downstream research outcomes and shows qualitative evidence of broader exploration. In a case study, it develops an extension that improves upon the reproduced baseline of an ICML 2026 oral paper excluded from our training corpus.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.