Monitoring and Adapting Deployed Decision-Focused Models with Sparse Feedback
Abstract
Decision-focused learning (DFL) trains predictive models to directly optimize downstream decision quality. Existing work addresses training and online updates with per-round feedback, but the deployment of a frozen DFL model under sparse feedback remains largely open. We study this setting: a trained predictor is deployed while the data distribution may drift, and the true cost vector is observed only through selective label queries subject to a finite budget. The operator must determine when decision quality has degraded and whether a candidate model deserves promotion. Yet regret is selectively observed, detection must remain valid at arbitrary stopping times, and an update triggered by an alarm can itself harm decisions. We propose E-REACT, a deployment framework whose detection rule provides finite-sample, anytime-valid false-alarm control and whose promotion gate controls erroneous deployments over the model's lifetime. We further characterize the optimal tradeoff between the label budget and detection delay, establish uniform sampling as the unique minimax allocation under volatility-reshaping shifts, and prove that scalar error monitors can be blind to regret drift. Across four combinatorial optimization problems and two real-world histories, E-REACT detects decision-relevant shifts missed by error monitoring. Its promotion gate blocks every spurious candidate under false alarms, and guarded updates reduce regret by up to 26.1% without harming any of 200 paired streams under genuine drift.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.