acceptodds
Under review as a conference paper at ICLR 2027

AI Research Preference Models

Abstract

AI research agents (AIRAs) can drive machine learning experiments from proposal to implementation and evaluation. Yet their progress on frontier tasks is often throttled by the cost of evaluation, which can consume days of GPU time. When an agent can propose far more candidates than it can afford to run, progress depends on its research preference: how it allocates a fixed execution budget across many candidates. We introduce AI Research Preference Models (RPMs) that predict which candidate solution is most promising, without needing to run them all. We build two variants of RPMs from pretrained language models: an inference-only model that reasons over candidate plans, code, and previously executed solutions, and an agentic model that also runs small-scale pilot experiments. Integrated into the AIRA-dojo research agent and evaluated on the machine learning research benchmark AIRS-Bench, the inference-only variant consistently provides gains over the unguided baseline with the agentic variant delivering the best overall performance. Across four LLM backbones (GPT-5, Gemma 4-31B, Qwen3.6-27B and Muse Glimmer), both variants improve the average normalized score over the baseline unguided agent. On average, they match the unguided agent's 24-hour performance in roughly 14 hours, using less than 60% of its execution budget, and they set new state-of-the-art results on three AIRS-Bench tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.