SchedulerFM: Reward-Free Pretraining for Generalizable Job-Shop Scheduling
Abstract
Recent studies using deep reinforcement learning (RL) for job-shop scheduling problems (JSSP) have largely relied on task-specific training, where separate policies are optimized for fixed scheduling objectives and settings. This limits reuse when the instance distribution, scale, objective, or scheduling setting changes. We propose SchedulerFM, a strategy foundation model for JSSP that decouples learning scheduling behavior from downstream optimization. SchedulerFM comprises two stages. First, reward-free RL pretraining learns a reusable latent strategy space by factorizing the long-horizon successor occupancies induced by diverse scheduling trajectories, without using a downstream scheduling reward. Second, latent-space test-time optimization freezes the pretrained model and searches over the strategy space, evaluating each candidate based on the downstream objective achieved by its complete rollout. A single JSSP-pretrained model transfers zero-shot across unseen distributions and scales, to flexible JSSP, and to previously unseen scheduling objectives, matching or outperforming strong learning and search baselines without additional training. More broadly, SchedulerFM opens a path toward reusable neural schedulers whose solution quality can scale with test-time compute without repeated task-specific retraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.