Learning Feature Interfaces to Steer Prescribed Contextual Bandits
Abstract
Many online decision systems are modular: an upstream component controls the features exposed to a downstream learner but cannot change the learner's reward or learning rule. While the learner optimizes its prescribed reward, the upstream component may optimize a different system-level objective. We study feature-interface optimization in repeated linear contextual-bandit deployments. Before each deployment, a leader applies a positive coordinate-wise rescaling to the raw features exposed to a prescribed contextual-bandit learner, which then learns from bandit feedback. Even this simple intervention leaves the follower's reward task unchanged while altering its finite-sample estimates, exploration, and collected feedback. As a result, the leader objective can depend nontrivially on the feature interface. We propose CILOT, which optimizes the interface using an unbiased trajectory-gradient estimator that differentiates through the follower's feedback-dependent online updates using only selected-action feedback. A single follower calibration guarantees regret uniformly over all admissible interfaces, while the leader optimization admits an projected-stationarity guarantee. Experiments on synthetic and real-world datasets demonstrate the benefit of optimizing the feature interface.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.