The Loss Is Only Half of the Step: Rethinking Decision Losses in Decision-Focused Learning
Abstract
Decision-focused learning replaces the squared error of a cost predictor with a decision loss and treats the choice of this loss as the main design decision. We find that the squared error already moves each prediction in a safe direction. Moving a prediction toward its realized cost never increases the regret of that sample when the objective is affine in the cost, and every task in our comparison has such an objective. Regret can rise only through the parameter update, which moves all predictions together and can carry a prediction off its path to the realized cost. Accordingly, we split a training step into the displacement that the loss assigns to each prediction and the transport by which one parameter update moves all predictions. Decision losses change only the displacement, so they correct a direction that is already safe. We introduce RIFT (regression in function space with trust regions), which fixes the transport as a damped Gauss-Newton step and accepts any displacement. RIFT with the regression displacement is not significantly worse than any of 22 methods on five tasks of PredictiveCO-Benchmark, and it trains 7 to 27 times faster than the strongest decision loss. Replacing the regression targets by the rescaled direction of the SPO+ surrogate or of a perturbed Fenchel-Young loss raises regret by 1.11 to 3.04 points on two knapsack tasks. Replacing the transport by gradient steps with a validated step size changes regret by at most 0.07 points, and both transports reach a regret 1.49 points below the regression baseline of the benchmark on energy prices. On these tasks, regression gains from a better transport and not from a decision loss.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.