UMAO: Coupling Cache Lifecycle and Admission for Agents on Unified Memory
Abstract
Multi-round agents continually append context, making cross-round reuse of computed prefixes essential for avoiding redundant prefill. On devices with a unified memory architecture (UMA), such as the NVIDIA DGX Spark, device and host share one physical memory pool. Expanding the host cache therefore does not increase total capacity; instead, it reduces the allocatable device KV, while retained histories, active KV, and cross-tier transfers compete for limited capacity. The system must maintain history usable in later rounds under fixed shared capacity and coordinate recovery and computation for competing requests. We propose unified memory agent offloading (UMAO), a framework that connects cache management and request admission through a shared recoverable-prefix state. Unified cache arena offloads long-term history to storage outside the shared pool, restricts host to bounded transfer staging, and commits and recovers history in continuous-prefix units. Cache-aware agent admission reads the prefix summary generated at the same scheduling boundary, jointly estimates direct-hit, cross-tier recovery, and recomputation work, and uses an age fence to limit tail latency. We implement UMAO as an SGLang HiCache plugin and evaluate it on DGX Spark with multiple real agent workloads. UMAO raises prefix reuse to 87.5%–90.3%, approaching the Ideal prefix-reuse upper bound across all workloads. Compared with existing HiCache write policies and longest-prefix match (LPM) admission, UMAO achieves up to / speedups in mean/P99 time to first token (TTFT) and / speedups in mean/P99 session completion time (JCT), respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.