When Predictions Become Supervision: Sharp Minimax Laws for Stochastic Replay
Abstract
When past predictors are reused to generate new training labels, each prediction shapes both current accuracy and future supervision. We study this feedback in online learning over thresholds with a fixed but unknown target. After an initial truthful round, an independent audit reveals the true label with probability , and otherwise an adversary-selected past predictor generates the feedback label. For horizons with fixed , we establish sharp minimax rates for every . At , the minimax cumulative loss is . Achieving this rate requires sacrificing some current accuracy to preserve future information, and this tradeoff is unavoidable even for improper learners. Restricting replay to sufficiently recent predictors on average reduces the loss to . For proper learners, revealing audit status or the selected replay source achieves the same rate. We further derive sharp rates for partial or noisy information about label generation and show that equally frequent disclosures can yield different minimax rates. Finally, with the replay source revealed but audit status hidden, fixing the reusable first predictor to raises the long-horizon loss from to .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.