The Memory Horizon Is the Lever: Benchmarking Online Active Learning under Concept Drift and Label Delay
Abstract
Active learning (AL) is studied almost entirely as an acquisition problem, and benchmarks on stationary data rank uncertainty sampling as the strategy to beat. In streaming deployment, however, purchased labels arrive after a delay and may describe a concept that has changed by the time they arrive. We study which decision of an online learner determines accuracy under these conditions: which points to query, or how to use the delayed labels once they arrive. We build a factorial benchmark that crosses eleven acquisition strategies with ten update rules on 50 audited temporal streams at three label delays, and propose an online horizon selector that chooses how many recent labels to learn from using only labels already paid for. No tested acquisition strategy improves on random labeling by more than 0.8 points in any cell at delays up to 500 steps, and a positive control shows that the gains acquisition achieves on stationary image features vanish once the relation between features and labels drifts. The update rule moves accuracy by 9 points, and its decisive property is the memory horizon: a buffer of the 50 most recent labels scores 77.5% against 67.9% for an unbounded one, yet loses up to 9 points on real streams whose drift is slow. The selector matches a per-stream oracle at delays up to 500 steps and costs 0.5 points where nothing drifts. Code and results: https://anonymous.4open.science/r/active-learning-cd-iclr27-1342/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.