acceptodds
Under review as a conference paper at ICLR 2027

RIGOR: Representation-Instrumented Generation with Outcome-Gated Refinement for Budgeted Discovery

Abstract

AI scientist systems built on large language models (LLMs) can generate hypotheses and design experiments, but evaluating a candidate can be costly in time, computation, or materials, so the choice of the next experiment matters. We present Representation-Instrumented Generation with Outcome-Gated Refinement (RIGOR), a framework for budgeted discovery that describes candidate designs with an explicit rubric of interpretable dimensions, which record how designs differ rather than how good they are. Measuring a candidate on these dimensions gives its rubric-based representation, which hypothesis generation and outcome prediction share. An LLM outcome predictor estimates each candidate's outcome and uncertainty from its complete design, its rubric-based representation, and previous outcomes, and an upper confidence bound (UCB) rule selects the next experiment. As evidence accumulates, RIGOR may propose new rubric dimensions, but it activates a dimension only if it reduces prediction error by at least 5% on the next five outcomes, with every forecast recorded before the outcome is observed. Across four scientific-design tasks and four MLS-Bench algorithm-design tasks, with Opus 4.7 and GPT-5.5, a budget of 24 evaluations, and three runs per setting, RIGOR outperforms existing baselines: it achieves the highest mean final best score in 11 of 16 model–task settings, the highest mean evaluated outcome in all 16, and the lowest within-run variation over the second half of the budget in 15. Its prospective gate admits only 16 of 86 proposed dimensions, each of which improves prediction of subsequent outcomes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.