Recovery Thresholds for Idealized Best-of-N Retraining
Abstract
Repeated best-of- distillation changes the model distribution over time, so rare errors may persist and the proportions of correct answers may drift. We study when mixing a fixed reference distribution into each retraining round guarantees recovery of the desired distribution. When all correct answers share the same score distribution, we give an exact necessary-and-sufficient condition for removing incorrect answers, based on a rare-error multiplier that measures whether selection amplifies a rare error. This shows that pairwise verifier accuracy alone does not determine long-run recovery. When correct answers have different score distributions, recovery further requires selection to preserve the reference distribution. Under this condition, we prove matching upper and lower bounds showing that, for fixed , the worst-case reference weight needed for global recovery scales as , where measures score-distribution differences. We also give sufficient conditions for the general heterogeneous setting, with exact or approximate recovery depending on preservation error.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.