Maia: Efficient LLM-based Agent Serving with Dynamic Resource Management
Abstract
Large Language Model (LLM)-based agents, customized through specific model adapters, system prompts, and external tools, are increasingly adopted to power intelligent services across different domains. In emerging Agent-as-a-Service platforms, service providers can offer multiple agent instances of different types on a pool of GPU resources. However, existing systems often rely on static agent management, resulting in poor elasticity and low efficiency when workloads fluctuate. We introduce Maia, an efficient agent serving system centered around the concept of lightweight agent reconfiguration, which enables dynamic switching of agent types on shared instances at runtime to meet dynamic agent workloads. Maia incorporates two key designs. First, it implements a fast agent reconfiguration mechanism that accelerates reconfiguration process via two techniques: (i) a consistency-aware prompt Key-Value (KV) reuse technique that reduces prompt initialization latency by reusing shared prompt segments, and (ii) a hybrid state KV migration strategy that overlaps state KV computation with communication to accelerate runtime state transfer. Second, we design an online agent management algorithm that dynamically reconfigures agent instances to balance serving throughput and reconfiguration cost. Extensive experiments using the defined utility loss metric show that Maia achieves up to a 52% reduction in utility loss compared to state-of-the-art baseline methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.