HexAGenT: Workflow- and Heterogeneity-Aware Scheduling for Agentic LLM Serving
Abstract
Serving agentic LLM applications requires scheduling dependent model calls to meet service-level objectives (SLOs) for complete user workflows. This is difficult when calls and branches emerge during execution, external tools delay progress, and prefill-decode (P-D) disaggregation places each call on separate, heterogeneous workers. Priorities based only on individual calls overlook workflow dependencies, while workflow priorities alone do not account for placement-dependent queueing, transfer, and KV-capacity costs. We present HexAGenT, a workflow- and heterogeneity-aware scheduler that couples workflow progress with joint P-D placement. HexAGenT maintains a critical-path horizon over revealed calls and completed tool phases, then normalizes projected call completion from workflow arrival by this horizon. It uses the resulting priorities to plan queue order and worker pairs, updating projected queues and capacity after each assignment. Local and global KV-prefix reuse refine per-worker prefill costs, and asynchronous plan application allows workers to continue serving during scheduling. In trace-driven simulation with two LLMs, four agentic workloads, and three homogeneous and heterogeneous cluster settings, HexAGenT reduces the SLO scaling factor required for and workflow attainment by up to and , respectively, relative to the best evaluated baseline at each target. Gains also persist on homogeneous workers, and additional experiments characterize sensitivity to service-time estimation error and aggregate planning cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.