Actor Score Composition and Refinement for Multi-Objective Flexible Job Shop Scheduling
Abstract
Multi-objective flexible job shop scheduling has attracted growing research interest due to the need to balance competing objectives in production. Current neural approaches use preferences to condition model features, parameters, or computation paths, enabling a single model to construct schedules across different trade-offs. However, generating high-quality schedule sets is challenging because a single model needs not only to learn assignment and sequencing rules required across preferences but also to continually adapt its decisions to changing machine workloads and remaining operations. To address this challenge, we propose Actor Composition with Preference Correction (AC-PC), which constructs preference-conditioned scheduling policies by composing and refining independent actor scores. Specifically, a shared conditional encoder represents the current scheduling state at fixed preference anchors, and independent actors produce action scores that are combined using weights determined by the requested preference. Moreover, preference correction uses the requested preference and mixed actor features to apply an action-specific adjustment to the composed scores. To evaluate AC-PC, experiments are conducted on same-size synthetic instances, larger unseen instances, and public benchmarks. The results show that AC-PC outperforms baselines in most cases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.