acceptodds
Under review as a conference paper at ICLR 2027

SAPO: Self-Anchored Preference Optimization for Cold-Start Neural Scheduling

Abstract

Constructive neural solvers for scheduling problems (e.g., JSSP, FJSP) face a critical cold-start dilemma: training from scratch lacks an informative reference prior, making standard preference optimization difficult to apply effectively. While recent reference-free methods alleviate the dependency on expert policies, they may exhibit two stage-specific weaknesses: premature convergence associated with deterministic winner selection in early training, and policy drift when no explicit trust region is available during late-stage refinement. To bridge this gap, we propose Self-Anchored Preference Optimization (SAPO), a stage-adaptive framework that dynamically constructs the missing prior from the training trajectory itself. SAPO introduces a dual-phase mechanism: (1) Distribution-Smoothed Exploration, which samples winners from a decaying top-K pool to broaden early winner supervision; and (2) Self-Anchored Refinement, which utilizes the historical-best policy as a moving anchor to provide an anchor-relative signal during refinement. Across six JSSP benchmark families, SAPO yields lower mean gaps than BOPO under a matched fixed-checkpoint decoding audit and shows improved stability on the recorded late-stage trajectory, while controlled FJSP/TSP studies provide preliminary evidence of transfer to related constructive tasks without systematic inference overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.