acceptodds
Under review as a conference paper at ICLR 2027

Delayed Heuristic-Behavior-Target Coupling for Off-Policy Reinforcement Learning in Inventory Control

Abstract

Many reinforcement learning methods tie data collection to the task policy, while coordinating independently optimized behavior and target policies remains challenging. Policies learned by these methods can perform unreliably in long-horizon tasks under uncertainty, such as inventory control, where failures can incur extreme costs. We introduce Heuristic Behavior Policy Optimization (HBPO), an off-policy framework that routes heuristic guidance through persistent behavior learning, coupled with target-policy, updating separately with residual-based advantage estimation and emphatic weighting. Under matched environment interaction budgets on DSIS and its risk-augmented variant DSIS-RB, HBPO recovers many high-cost initial policies and improves most non-collapse comparison groups. Cold starts achieve competitive costs without final cost collapse across all tested seeds. We identify a transmission chain in the delayed heuristic–behavior–target coupling as a key driver of recovery and improvement. Generation-level diagnostics trace heuristic uptake and subsequent selective target responses, while component ablations support the contribution of this transmission mechanism to policy recovery and improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.