SAPO: Self-Anchored Preference Optimization for Cold-Start Neural Scheduling
Abstract
Constructive neural solvers for scheduling problems (e.g., JSSP, FJSP) face a critical cold-start dilemma: training from scratch lacks an informative reference prior, making standard preference optimization difficult to apply effectively. While recent reference-free methods alleviate the dependency on expert policies, they may exhibit two stage-specific weaknesses: premature convergence associated with deterministic winner selection in early training, and policy drift when no explicit trust region is available during late-stage refinement. To bridge this gap, we propose Self-Anchored Preference Optimization (SAPO), a stage-adaptive framework that dynamically constructs the missing prior from the training trajectory itself. SAPO introduces a dual-phase mechanism: (1) Distribution-Smoothed Exploration, which samples winners from a decaying top-K pool to broaden early winner supervision; and (2) Self-Anchored Refinement, which utilizes the historical-best policy as a moving anchor to provide an anchor-relative signal during refinement. Across six JSSP benchmark families, SAPO yields lower mean gaps than BOPO under a matched fixed-checkpoint decoding audit and shows improved stability on the recorded late-stage trajectory, while controlled FJSP/TSP studies provide preliminary evidence of transfer to related constructive tasks without systematic inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.