What Early Stopping Hides: Cost-Aware Auditing in Post-Training
Abstract
Early stopping saves computation when selecting post-training updates but leaves the outcomes of rejected runs unknown. We introduce ReSight, which uses predictions from intermediate training results and remaining costs to sample stopped runs for completion and evaluation. Each completed run is compared with the original choice using constraints and costs fixed at screening. We correct for unequal sampling probabilities and count unresolved comparisons as possible errors to estimate the mean number of missed improvements per decision. Under stated sampling and evaluation assumptions, we derive upper bounds on the true mean over a fixed set of decisions. On the Berkeley Function Calling Leaderboard (BFCL) static benchmark with Qwen2.5-1.5B-Instruct, ReSight reduces the root mean squared error by 28.8% for decisions that retained the original model, compared with selecting runs based only on their remaining costs, for the same expected expenditure. In a prospective audit of 120 model selection decisions, estimates are close to reference counts, while measured cost, including setup, is 24.0% lower than the cost of completing and evaluating all runs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.