acceptodds
Under review as a conference paper at ICLR 2027

BIP: Online Scheduling for LLM Batch Inference with Predictions

Abstract

The growing deployment of large language models (LLMs) has created a large demand for offline batch inference, where jobs must be executed under deadlines on heterogeneous GPU clusters. Unlike online serving, this setting is throughput-driven and job-oriented. We introduce BIP, a scheduling framework that combines output-length prediction, rolling-horizon integer programming, and online correction of underestimated workloads. The scheduling model captures job deadlines, GPU-type-dependent speeds, parallel execution, and slot-boundary preemption. We develop three formulations for completed jobs, normalized service, and completion time, and establish the interpretation and feasibility conditions of their objectives. Simulations using batches constructed from instruction–response data show that the best-of-three formulation envelope improves deadline-completion rates over FCFS and SJF across the tested configurations. At normalized load 1.0 and relative job size 1.0, the corrected envelope completes 55.3% more jobs than FCFS and 32.4% more than SJF. These results demonstrate the potential of combining workload predictions with job-level deadline scheduling for offline LLM inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.