acceptodds
Under review as a conference paper at ICLR 2027

Learning Advantage Function Surrogates via Convex Conjugate Matching

Abstract

A policy trained to make per-unit decisions, such as how much of a product to order, tells us the best action but not the cost of deviating from it. That cost, the advantage function, is what downstream systems need when they have to depart from the per-unit optimum: to estimate the impact of ordering more or fewer units, to respond to a one-off vendor discount, or to satisfy a shared capacity constraint. The advantage function is what makes such trade-offs computable, yet pure policy-learning algorithms often produce the policy alone, without a critic that would yield it. We propose a method to train an advantage function surrogate through price perturbations motivated by a Fenchel-Young loss. We describe a two-stage recipe (fit the policy, then the cost model) and illustrate it on a multi-period inventory management problem.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.