When to Commit and When to Defer: Maturing Markov Decision Processes under Refining Information and Expiring Opportunities
Abstract
Sequential decisions often become better informed as action opportunities expire. We introduce Maturing Markov Decision Processes (MMDPs), a structured subclass of augmented finite-horizon Markov decision processes, for organizing when to commit and when to defer. The Expiring-Actions-First Principle provides a sufficient condition for deferral based on information gain and delay loss. This structure motivates stage-local policies, action abstraction, and policy-guided planning. Replenishment experiments show improved learning and test performance under the MMDP interface and, in the primary setting, a larger test-performance gain from forecast refinement. In synthetic cash management, MMDP improves PPO test returns across both network sizes; five-account equal-size mask controls show that action count alone does not explain its learned gain, and the advantage persists with action abstraction and policy-guided search. A production-scale cash-management case study adds complementary evidence. Together, these results support commitment timing as a useful inductive bias for learning and planning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.