Stage-Adaptive Submodular Objective Selection for Data-Efficient Training
Abstract
The most useful data-selection objective can change as a model learns, yet existing bandit-based selectors rely on historical rewards that can lag the model's current needs. We introduce DriftBandit, which scores candidate objectives before updating the model using gradient-validation alignment. By comparing candidate-subset gradients with a held-out validation gradient, it estimates objective utility at the current parameters. A complementary Page-Hinkley trigger temporarily increases exploration after validation deterioration. An arm-sampling variant reduces per-decision scoring cost from O(K) to O(m) by evaluating only m of K objectives. Across three vision and six medical-imaging benchmarks, each at three per-batch update budgets, full-information DriftBandit outperforms OnlineSubmod in all 27 settings. It also surpasses sample-level dynamic selectors and non-stationary bandits in the vision experiments, demonstrating gains beyond the closest objective-selection baseline. Fixed-objective comparisons establish the benefit of adapting the selection criterion during training, while ablations identify immediate gradient alignment as the principal source of improvement and deterioration-triggered exploration as a complementary gain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.