acceptodds
Under review as a conference paper at ICLR 2027

Policy Capacity and System Cost in Independent Reinforcement Learning for Serial Inventory Control

Abstract

Increasing the expressiveness of a policy can change how independently trained agents affect one another. We investigate this effect in a serial inventory system in which each agent observes local state and minimizes its own quadratic inventory cost. Under a fixed independent proximal policy optimization protocol, we compare a one-gain feedback policy, an affine policy, and a residual multilayer perceptron. In a four-echelon chain with lead time two and demand autocorrelation 0.6, mean system cost increases from 40.679 for the one-gain policy to 55.379 for the affine policy and 56.194 for the neural policy, corresponding to increases of 36.1% and 38.1%. Every restricted-policy run is cheaper than every richer-policy run across eight training seeds per class. Paired bootstrap intervals for the absolute increases are [13.57, 15.78] and [14.03, 17.21]. Across additional demand regimes, the affine cost gap and sensitivity to recent demand increase together, but changing marginal demand variance prevents a causal interpretation. We also prove that increasing a stable proportional feedback gain raises upstream variance costs in an idealized linear cascade. The theorem provides an externality argument, while the experiments establish a finite-budget learning effect; they do not certify Nash convergence or isolate a causal neural mechanism. These findings support evaluating policy classes by their system-level consequences when learning objectives are decentralized.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.