Online Learning with Improving Proxies: A Multi-Fidelity Bandit Approach
Abstract
In many online learning problems, reliable evaluation can be costly, motivating inexpensive proxy feedback that approximates a trusted target. While proxy quality is often treated as static, learned evaluators may improve through continued training, self-generated supervision, or computational refinement, creating a trade-off between further proxy investment and trusted feedback. Motivated by such self-evolving evaluators, including LLM-based judges, we study a multi-fidelity multi-armed bandit with a stationary high-fidelity target and a low-fidelity source whose discrepancy evolves with use. We introduce an a priori bound on the average proxy-target mismatch, derive improvement-aware confidence bounds, and propose the Threshold-Based Adaptive Continuation Companion (TACC), which determines whether to continue low-fidelity evaluation or escalate to high fidelity. We establish an instance-dependent regret bound showing that limited low-fidelity continuation eliminates the logarithmic high-fidelity term for a class of suboptimal arms, and a lower bound capturing the cost of proxy improvement. Experiments on synthetic bandits and NAS-Bench-201 evaluate when adaptive continuation improves performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.