Learning to Construct and Improve for Large-Scale Job Shop Scheduling via Deep Reinforcement Learning
Abstract
Learning-based Job Shop Scheduling Problem (JSSP) solvers have made substantial progress through graph representations, size-generalizable policies, and learned improvement. Yet learning directly on large scheduling instances remains challenging: the solver must represent complex precedence and resource dependencies over long horizons while efficiently allocating computation across large decision and search spaces. We study how to make large-scale JSSP instances first-class training states and propose Learning to Construct and Improve (L2CI), a deep reinforcement learning framework that jointly learns schedule construction and improvement at scale. For scalable representation learning, we develop a structure-aware self-attention Transformer that captures job and machine dependencies, while inducing-point compression provides fixed-size global representations. During construction, successor-state bounds supervise soft pruning of unpromising actions. After construction, an RL-guided critical-path search improves complete schedules and recycles improved solutions as demonstrations for the constructive policy. We train L2CI directly on instances with up to 12,000 operations and achieve strong performance on standard TAI and DMU benchmarks, with learned search further improving the constructed schedules.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.