Rehearse: Stepping Back from the Confidence Cliff in Autoresearch
Abstract
Improving a machine-learning system and maintaining the ability to identify its next improvement are distinct challenges for autoresearch. We isolate pre-execution judgment from proposal generation and execution, and examine how it changes as successful modifications accumulate. We construct a 366-pair benchmark from 39 paper-derived tasks in public AutoSOTA logs (Li et al., 2026; Tsinghua FIB Lab, 2026). Its primary 296 same-baseline pairs each match one modification that produced an accepted improvement with one that did not. With measured outcomes hidden, an LLM judge given candidate rationales but no prior-attempt history becomes more consistent across presentation orders but less accurate as successful changes accumulate. On the full benchmark, accuracy on agreed verdicts (selective accuracy) falls from 82.8% to 56.9%, while the fraction of pairs receiving an agreed verdict (strict-consensus coverage) rises from 76% to 85%. We call this mismatch the *confidence cliff*. **Rehearse** is a lightweight skill that proposes several ideas, retrieves similar past attempts and their outcomes to compare the candidates, runs the most promising, and records its outcome for future judgments. This focused outcome memory raises late selective accuracy to 83.5%. Across budgeted training runs over three loops, Rehearse improves the endpoint under the same training-run budget on nanochat, image classification, and time-series forecasting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.