Don't Learn When to Trigger, Learn When Waiting Is Not Worthwhile: A New Framework for Dynamic Task–Resource Assignment
Abstract
For large-scale task–resource assignment problems in dynamic environments with hard constraints, a common solution is a hierarchical RL–heuristic framework, in which an upper-layer reinforcement learning (RL) policy determines which tasks to trigger at each decision epoch, while a lower-layer heuristic assigns available resources to them. However, since most pending tasks remain untriggered at most decision epochs, RL receives extremely sparse reward signals, leading to ambiguous credit assignment and ineffective triggering policies. Our key insight is that, instead of asking RL to repeatedly decide whether to trigger at each decision epoch, we can reformulate the problem as determining how much can still be gained from waiting. A task is triggered when its delay benefit, defined as the expected-return difference between its best triggering opportunity in the remaining life cycle and triggering it now, falls below a learned threshold. Based on this insight, we propose a new triggering framework that decomposes the original triggering problem into two complementary learning problems. First, a delay-benefit estimator is trained with rollout-generated supervision to estimate how much can still be gained by waiting. Second, RL selects a task-specific threshold that determines when the estimated delay benefit is sufficiently small to trigger the task, while accounting for global competition among tasks and resources. This decomposition replaces repeated RL triggering decisions with supervised delay-benefit estimation and one-time RL threshold selection, providing dense supervision for delay-benefit estimation while allowing RL to focus on long-horizon coordination. Experiments on Dynamic Job-Shop Scheduling, Weapon-Target Assignment, and Warehouse Replenishment demonstrate consistent improvements over pure RL and hybrid RL–heuristic approaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.