acceptodds
Under review as a conference paper at ICLR 2027

SchedulerFM: Reward-Free Pretraining for Generalizable Job-Shop Scheduling

Abstract

Recent studies using deep reinforcement learning (RL) for job-shop scheduling problems (JSSP) have largely relied on task-specific training, where separate policies are optimized for fixed scheduling objectives and settings. This limits reuse when the instance distribution, scale, objective, or scheduling setting changes. We propose SchedulerFM, a strategy foundation model for JSSP that decouples learning scheduling behavior from downstream optimization. SchedulerFM comprises two stages. First, reward-free RL pretraining learns a reusable latent strategy space by factorizing the long-horizon successor occupancies induced by diverse scheduling trajectories, without using a downstream scheduling reward. Second, latent-space test-time optimization freezes the pretrained model and searches over the strategy space, evaluating each candidate based on the downstream objective achieved by its complete rollout. A single JSSP-pretrained model transfers zero-shot across unseen distributions and scales, to flexible JSSP, and to previously unseen scheduling objectives, matching or outperforming strong learning and search baselines without additional training. More broadly, SchedulerFM opens a path toward reusable neural schedulers whose solution quality can scale with test-time compute without repeated task-specific retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.