COMAS: Co-Optimizing Model and Policy in Stochastic Contextual Bandits
Abstract
The performance of contextual bandit algorithms critically depends on the model configuration used to estimate the reward function, including its family, architecture, and regularization. However, selecting an effective configuration requires data collected by the action-selection policy, which itself depends on the reward estimator, creating a circular dependency between model configuration and policy learning. We introduce COMAS, a framework that co-optimizes model configuration and action selection by treating configuration as a sequential decision variable. At each episode, COMAS selects a model configuration, trains the corresponding reward estimator on historical interaction data, and uses a Gaussian process surrogate to model the episode-level reward as a function of the configuration, thereby guiding subsequent configuration selection. Under suitable assumptions, we show that COMAS achieves a sub-linear regret upper bound that separately accounts for regret due to configuration selection and policy learning. Experiments on synthetic and real-world contextual bandit problems demonstrate the effectiveness of co-optimizing heterogeneous model configurations and action-selection policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.