acceptodds
Under review as a conference paper at ICLR 2027

When Do Optimization Duals Help Reinforcement Learning?

Abstract

In reinforcement learning with embedded optimization, inner solvers return Karush–Kuhn–Tucker (KKT) dual variables that provide structured abstractions of constraint and resource pressure induced by the current decision context, yet how such representations affect actor–critic learning remains unclear. We characterize when such representations improve critic learning and, when action-dependent, how they affect actor-gradient identifiability. Since the solver representation is a deterministic function of the complete optimization context , explicitly providing it to the critic adds no Bayes information. Nevertheless, the added representation can expose target-relevant solver structure that is difficult for the baseline class to recover. For nested linear or frozen-feature classes, we derive an exact fixed-design finite-sample risk decomposition balancing captured structural signal against added estimation cost, with a separate Gaussian reference crossover and a bridge to fixed-policy value learning. On the actor side, we show that the detached action partial derivative is generally not identified by on-graph critic values, with its ambiguity confined to . Matched total-derivative and score-function interfaces are instead extension-invariant. Experiments spanning controlled diagnostics, fixed policy value learning, and full actor–critic control support the predicted critic-side learnability trade-offs and actor-side gradient-identifiability effects.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.