When Fixed Thresholds Fail: Trace-Grounded Scheduling for Multi-Tenant LLM Fine-Tuning under SLOs
Abstract
Multi-tenant platforms can schedule LLM fine-tuning by watching each job's evaluation curve, terminating runs whose score declines and handing the freed GPUs to another tenant. A recently published rule of this kind reports a false-positive rate. We replay it on real per-checkpoint trajectories from public pretraining runs. Every run is healthy, so every termination is an error, and the rule fires on of them. Neither obvious explanation holds up. The synthetic concave curve used for calibration predicts an even higher rate, and real evaluation noise is milder than assumed rather than harsher. What actually carries the published number is a quantity it never states, the per-interval signal-to-noise ratio. Real curves sit at a median of , whereas at demands . Correct that single parameter and the original generator reproduces our measurement to within four points. Raising does not rescue the rule, because real declines cluster, and a classical quickest-detection bound puts the delay needed for a false-alarm rate at three times the checkpoints a job has, so by that bound no causal detector reaches the published operating point on these curves. Standard reporting would not have caught this. Replaying a -job production trace, we find that how SLO attainment is aggregated moves the reported number further than which scheduler produced it, by as much as points, and reorders the policies outright. Early termination then corrupts the usual metrics in one direction. On a workload with no declining jobs at all, the published rule “improves” mean job completion time by while destroying of useful compute. At the trace's actual prevalence of decline, nothing we tested pays for itself, including a learned policy that never beats a one-parameter rule normalising drawdown by an online noise estimate. We recommend net goodput reported with a false-positive rate quoted at its checkpoint cadence, the one pairing that survived every probe we ran.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.