acceptodds
Under review as a conference paper at ICLR 2027

Rehearse: Stepping Back from the Confidence Cliff in Autoresearch

Abstract

Improving a machine-learning system and maintaining the ability to identify its next improvement are distinct challenges for autoresearch. We isolate pre-execution judgment from proposal generation and execution, and examine how it changes as successful modifications accumulate. We construct a 366-pair benchmark from 39 paper-derived tasks in public AutoSOTA logs (Li et al., 2026; Tsinghua FIB Lab, 2026). Its primary 296 same-baseline pairs each match one modification that produced an accepted improvement with one that did not. With measured outcomes hidden, an LLM judge given candidate rationales but no prior-attempt history becomes more consistent across presentation orders but less accurate as successful changes accumulate. On the full benchmark, accuracy on agreed verdicts (selective accuracy) falls from 82.8% to 56.9%, while the fraction of pairs receiving an agreed verdict (strict-consensus coverage) rises from 76% to 85%. We call this mismatch the *confidence cliff*. **Rehearse** is a lightweight skill that proposes several ideas, retrieves similar past attempts and their outcomes to compare the candidates, runs the most promising, and records its outcome for future judgments. This focused outcome memory raises late selective accuracy to 83.5%. Across budgeted training runs over three loops, Rehearse improves the endpoint under the same training-run budget on nanochat, image classification, and time-series forecasting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.