CausaSched: A Causal-Guided Deep Reinforcement Learning Hyper-Heuristic for Improving Multiple Shop Scheduling Problems
Abstract
Traditional constructive heuristics and most learning-based scheduling methods generate schedules from scratch, whereas recent learning-based improvement methods start from existing feasible solutions and further improve their quality through successive neighborhood moves. Because schedule improvement concerns how changes in controllable decisions affect the makespan, this process can be modeled from a causal-intervention perspective. Based on this observation, we introduce causal reasoning into learning-based hyper-heuristics and propose CausaSched for multiple scheduling problems. Causal reasoning starts from scheduling bottlenecks, traces how delays arise through structural dependencies, and determines where subsequent schedule modifications should be applied. Achieving this goal raises three key questions: how to efficiently search for intervention targets that can reduce makespan, how to match them with neighborhood operators satisfying problem-specific constraints, and how to jointly train both decision networks to improve search performance. CausaSched exploits job-precedence and machine-order dependencies shared across scheduling problems: a causal reasoning network learns where to intervene, and an operator selection network determines the neighborhood move. We further propose Causal Hierarchical Group Relative Policy Optimization (CH-GRPO), which jointly trains both networks using multiple multi-step search trajectories and expands the neighborhoods of high-quality states. Experiments on public benchmark instances and mixed synthetic instances show that CausaSched significantly outperforms deep reinforcement learning and improvement heuristics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.