acceptodds
Under review as a conference paper at ICLR 2027

Fusing Optimization Proxies and Reinforcement Learning for Multistage Stochastic Optimization

Abstract

Multistage stochastic programs provide a general framework for sequential decision-making under uncertainty, but their complexity grows exponentially with the number of decision epochs. A popular policy for multistage stochastic optimization solves a two-stage stochastic program at every time step in a rolling-horizon manner, relaxing non-anticipatory constraints for subsequent stages. Solving this sequence of optimization models is computationally expensive, limiting real-time deployment for some classes of applications. This paper presents PROSE (Proxies Reinforcement-learned in Optimization-Structured Environments), which fuses optimization proxies with deep reinforcement learning to obtain a feasible policy for the multistage problem. The optimization model defines the learning problem: at each epoch, the policy takes the same inputs as the two-stage program, including its weighted scenarios. Its action space is the feasible set of the optimization model, enforced by a differentiable repair layer applied to a neural prediction, and its reward is the first-stage cost. The policy is trained by backpropagating the cumulative realized cost through repair and state transitions over complete episodes, end-to-end in a self-supervised manner. On energy dispatch and inventory replenishment case studies, PROSE implements a feasible action at every decision, is 1,600 faster per decision, and typically incurs a lower cumulative realized cost than the rolling two-stage policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.