acceptodds
Under review as a conference paper at ICLR 2027

Scheduling Kernel-Generating Agents at Scale in Decoupled Agent–Worker Systems

Abstract

Programming agents develop GPU kernels by repeatedly generating, compiling, and running code. In production, the large volume of operator-generation tasks motivates a decoupled Agent–Worker architecture that avoids dedicating GPUs to individual agents during model inference and tool use. Agents run on CPU clusters and submit code under test to a shared GPU execution pool when compilation and testing are needed. In this setting, each batch’s Skill, the types and number of operators to generate, and model inference speed jointly determine how frequently the batch requires GPU execution. Too few execution requests leave GPUs idle, whereas too many cause queues to accumulate, delaying execution feedback and reducing useful iterations, which can compromise kernel quality. We propose a Scheduling Agent that controls the concurrency and admission of operator-generation tasks. It reads each batch’s Skill to identify the stages that require GPUs, then combines resource utilization, the queue of code awaiting execution, and model-API feedback to select the next tasks to launch. It also uses file semantics to promptly identify and remove residual artifacts created at incorrect output paths, reclaiming disk space and inodes. Compared with FIFO alone, adding the Scheduling Agent increases Worker utilization from 59.9% to 85.2%. It also removes approximately 2.2 TB of residual files from incorrect paths and frees about 380,000 inodes. The system has operated in production for several months, providing scheduling support for over ten thousand operator-generation tasks each month.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.