acceptodds
Under review as a conference paper at ICLR 2027

IGT-OMD: Implicit Gradient Transport for Decision-Focused Learning under Delayed Feedback

Abstract

Decision-focused learning (DFL) trains prediction models for better decisions, rather than prediction accuracy alone. In the online setting, predictions can guide successive decisions; a learning algorithm (the learner) updates a model as decision outcomes arrive. When those outcomes arrive late, both the model and the decisions it would produce have changed, making old learning gradients stale. How can the learner make better use of this feedback? We introduce implicit gradient transport for online mirror descent (IGT-OMD) to answer this question and combine gradients from new feedback with changes in past gradients caused by model updates. One version of this algorithm reuses stored decisions and gradient information (frozen transport); another recomputes them at the current model (full re-evaluation). For stationary problems, we bound the frozen gradient-descent variant's excess cumulative loss over the best fixed model chosen in hindsight (i.e., regret), and derive a sufficient number of iterations to solve the decision problem. For a shared strongly convex quadratic model, continuous and discrete stability analyses give an explicit learning-rate rule. Experiments cover linear–quadratic regulator (LQR) stability, Warcraft paths, and Sinkhorn transport. At delay 50 on Sinkhorn, full re-evaluation with Adam reduces loss by 10.81% versus learning from newly arrived feedback alone and by 26.6% versus the best tested delayed optimistic mirror descent (DOOMD) baseline. The transport gain costs the arrival-only control's training-update time. Independent random realizations confirm the transport benefit; frozen transport offers a lower-cost option in the tradeoff between decision quality and computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.