Align-and-Improve: Addressing State Staleness and Plan Quality in Dual-Stream Dynamic Scheduling
Abstract
Dynamic job shop scheduling requires immediate assignments, while slower search constructs candidate rules and plans. In a dual-stream scheduler, a request may be based on an observation that becomes stale before generation starts. A returned rule may also produce an inferior continuation, and an accepted plan may lose its advantage as execution proceeds. We present Align-and-Improve (AAI), which addresses these distinct sources of mismatch at request, plan and action time. Late State Binding (LSB) associates a pending request with the latest observed state when the generator claims it. Verified Plan Update (VPU) materializes candidates on the current state and replaces an incumbent only after a feasible, strictly improving comparison and a commitment-time consistency check. Rollout-Based Action Arbitration (RAA) compares the retained plan's next action with an action from the fast dispatching rule, evaluating both with the same fast-policy continuation before executing one assignment. Across 58 instances from four benchmark families, AAI lowers mean relative percentage deviation by about five percentage points versus RACE on three shared Qwen3.5 backbones while cutting model-token usage by up to about 90%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.