acceptodds
Under review as a conference paper at ICLR 2027

When Predictions Become Supervision: Sharp Minimax Laws for Stochastic Replay

Abstract

When past predictors are reused to generate new training labels, each prediction shapes both current accuracy and future supervision. We study this feedback in online learning over thresholds with a fixed but unknown target. After an initial truthful round, an independent audit reveals the true label with probability , and otherwise an adversary-selected past predictor generates the feedback label. For horizons with fixed , we establish sharp minimax rates for every . At , the minimax cumulative loss is . Achieving this rate requires sacrificing some current accuracy to preserve future information, and this tradeoff is unavoidable even for improper learners. Restricting replay to sufficiently recent predictors on average reduces the loss to . For proper learners, revealing audit status or the selected replay source achieves the same rate. We further derive sharp rates for partial or noisy information about label generation and show that equally frequent disclosures can yield different minimax rates. Finally, with the replay source revealed but audit status hidden, fixing the reusable first predictor to raises the long-horizon loss from to .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.