acceptodds
Under review as a conference paper at ICLR 2027

One State Variable Was the Whole Effect: A Controlled Study of Offline-to-Online Reinforcement Learning for Cluster Scheduling

Abstract

Pretraining a cluster scheduler on the production scheduler's placement log, fine-tuning online, and protecting the learned policy with a safety layer and an out-of-distribution fallback is a common recipe, yet its offline premise is rarely tested. We evaluate this pipeline on a Google production trace replayed in a capacity-constrained simulator. The offline premise fails: production placements show little predictable dependence on system load and cannot be reliably recovered from any shortlist we can construct. The deployed -slot shortlist recalls of them, and a -candidate shortlist containing every machine a job has ever used reaches only , making behaviour cloning the poorest-performing model we train. Two implementation defects initially obscured this failure: using only one gradient step per update and a policy head degenerate at initialisation. The dispersion of candidate logits, rather than policy loss, reveals both defects. After correcting the pipeline, the learned policy outperforms the production scheduler. However, its apparent advantage over heuristic baselines disappears when those baselines are given the co-runner count, a state variable omitted from the original feature set: adding this variable shifts the performance frontier by 5.61 percentage points and reduces the learned policy's apparent wins from of operating points to of , while denser sampling within the same heuristic family does not change the result. We therefore argue that baseline feature coverage should be established before attributing performance gains to learning. Replication on two additional traces, including one from a different operator, yields the same qualitative finding while showing that the magnitude of the gain depends strongly on simulator capacity: the gain is 0.29 percentage points at the trace's native cluster size but 43.3 percentage points in the capacity-shrunk setting. Thus, the observed improvement should not be interpreted as a general ranking of the production scheduler. Finally, a deployable takeover-and-fallback gate reduces catastrophic episodes from of to of . The source code and the result artifacts behind every number reported here are provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.