Prediction Losses and the Value of Hints in Optimistic Mirror Descent
Abstract
Prediction-dependent regret bounds motivate training losses for online decision systems, but statistical calibration of a regret remainder does not ensure task improvement. We study this gap for fixed-step optimistic mirror descent with full feedback and exogenous costs. The exact positive remainder is a Bregman loss calibrated to the conditional cost mean; subtracting the prediction-dependent stability term recovers task loss. We characterize policy differences through inverse-Hessian sensitivity. At any polytope vertex exposed by a cumulative linear margin, including nonsimple vertices, all bounded hint policies have absolutely convergent loss differences and share a logarithmic regret coefficient equal to the total active barrier weight divided by the step size. On a simplex approaching a face, the controller instead retains a lower-dimensional prediction problem, with a uniform logarithmic cumulative approximation error. A paired-cost construction then yields a linear task advantage for a learned saturated policy over a learned mean predictor, even as the latter's excess remainder tends to zero. We also establish a scalar boundary-curvature transition and matching estimation bounds under shrinking prediction geometry. Together, these results identify when calibration is statistically sufficient and when the controller's surviving decision directions make the prediction target consequential.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.