acceptodds
Under review as a conference paper at ICLR 2027

Learning Feature Interfaces to Steer Prescribed Contextual Bandits

Abstract

Many online decision systems are modular: an upstream component controls the features exposed to a downstream learner but cannot change the learner's reward or learning rule. While the learner optimizes its prescribed reward, the upstream component may optimize a different system-level objective. We study feature-interface optimization in repeated linear contextual-bandit deployments. Before each deployment, a leader applies a positive coordinate-wise rescaling to the raw features exposed to a prescribed contextual-bandit learner, which then learns from bandit feedback. Even this simple intervention leaves the follower's reward task unchanged while altering its finite-sample estimates, exploration, and collected feedback. As a result, the leader objective can depend nontrivially on the feature interface. We propose CILOT, which optimizes the interface using an unbiased trajectory-gradient estimator that differentiates through the follower's feedback-dependent online updates using only selected-action feedback. A single follower calibration guarantees regret uniformly over all admissible interfaces, while the leader optimization admits an projected-stationarity guarantee. Experiments on synthetic and real-world datasets demonstrate the benefit of optimizing the feature interface.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.