CARE: Label-Free Cost-Aware Online Routing for Zero-Shot VLM Classification
Abstract
Contrastive vision–language models (VLMs) exhibit complementary strengths and widely different computational costs in zero-shot classification, making a single fixed model suboptimal across diverse inputs. We introduce CARE, a label-free online framework for per-sample routing among heterogeneous CLIP-style VLMs under a global computational budget. Given a stream of unlabeled inputs, CARE dynamically selects one expert for each sample without requiring labeled calibration data or a pre-deployment reference set. To enable cross-model comparison, CARE standardizes logits on a per-sample basis and uses the resulting confidence as a label-free proxy reward. Routing is formulated as a contextual bandit with a dynamic dual update that balances estimated utility and computational cost while enforcing the global budget. Experiments across eight zero-shot classification benchmarks show that CARE achieves strong accuracy–cost trade-offs compared with individual models, full-expert methods, and routing baselines that use additional calibration or reference data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.