acceptodds
Under review as a conference paper at ICLR 2027

Priced Guidance: Can Language Models Generate Future Research Ideas?

Abstract

We evaluate language models' capability to generate novel research ideas through the lens of compression. Our goal is to lower-bound the potentially tiny probability that a language model generates the essence of a future research idea without any hints. Rather than estimate this probability through expensive repeated sampling, our Priced Guidance framework measures the compression cost: how many additional bits of information are needed to guide the model to recover the target idea. We prove that if the model can recover the target idea with at most bits of guidance in expectation, then it can generate the idea without any guidance with probability at least . In our framework, the language model, called the generator, can pose a sequence of multiple-choice questions, e.g., “what is the topic of the research idea?”, and specify a probability distribution over possible answers. A guide, which is a language model with access to the target idea, selects answers. If a selected answer has prior probability , the generator pays bits. The generator's goal is to minimize the cumulative cost it needs to pay to produce an idea that matches the essence of the target idea. This cumulative cost equals to, up to a constant, the number of bits of information sent by the guide. Using this methodology, we evaluate five generator models (Opus 5, Fable 5.1, GPT-6 Astra, GPT-5.6 Sol, and GLM 5.3) on the core ideas in 87 recent high-quality deep learning papers. We use a frontier model as a judge to determine whether the generated idea matches the target in terms of the central research object and defining mechanism. Fable 5.1 achieves the lowest median compression cost at 67.8 bits, substantially lower than gzip's median of 5,712 bits for losslessly compressing the summary of the target idea. A uniform ensemble of Fable 5.1, Opus 5, and GPT-6 Astra further reduces the median compression cost to 53.8 bits and improves the generation probability lower-bound by .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.