acceptodds
Under review as a conference paper at ICLR 2027

Returned-Predictor Audit for Early Model Selection

Abstract

Early model selection briefly trains several candidates, continues training on a subset, and uses validation to select a saved state for prediction. An improvement over random choices can come from favoring stronger candidates overall (usage), or assigning them to contexts where they help more (matching). We introduce Returned-Predictor Audit (RPA), an offline evaluation that separates these gains and traces them to the states actually returned. From complete records for the compared training plans, RPA evaluates random assignment at the policy's frequencies, computes the best reassignment at those frequencies using recorded test losses, and varies final validation while keeping early choices fixed. In an external two-task partial differential equation (PDE) study, a preselected pair yields the same states as the six-model pool in all 120 cases, with 41% less reconstructed training and validation work. In a six-task PDE study, positive usage and negative matching estimates nearly cancel each other out. On LCBench, pairing one fixed default with a second model chosen for each dataset reduces classification error by 0.250 percentage points compared to random assignment at the same frequencies. Broader final selection changes both measured matching and the opportunity left among returned predictors: matching can decrease while the policy's loss gap to the best reassignment under the new rule narrows.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.