Controlling Scale Mismatch in Decision-Focused Learning with Cost Regularization
Abstract
Decision-focused learning (DFL) trains predictive models to minimize the cost of the decisions they induce rather than prediction error. When the model predicts the objective coefficients of a combinatorial optimization problem, the map from costs to decisions is piecewise constant, and many state-of-the-art methods obtain informative gradients by perturbing the predicted costs. We show that these methods share a common weakness: the perturbation scale is fixed, but the scale of the predicted costs is not. Because decisions are unchanged when costs are rescaled, the training loss does not constrain that scale, and it can drift freely during training. When predicted costs grow too large relative to the perturbations, the perturbed decisions stop differing from the unperturbed ones, and learning stalls or reduces to imitation. When they become too small, the perturbations dominate and the gradients no longer reflect the prediction. Since this drift varies across samples and over training, tuning the perturbation scale alone cannot prevent it. We propose a simple, method-agnostic remedy: a smooth shrinkage of the predicted costs, which leaves decisions unchanged and adds negligible cost. Experiments with three DFL methods show that it markedly reduces regret on a problem designed to induce this instability, and on standard benchmarks it corrects severe training failures at little cost when none occur.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.