acceptodds
Under review as a conference paper at ICLR 2027

Model as Inference Advisor: Leveraging LLM Semantics for Agentic KV Cache Scheduling

Abstract

Agentic workloads powered by Large Language Models (LLMs) are becoming increasingly critical; however, their long contexts and concurrent executions generate massive key-value (KV) caches, imposing severe GPU memory pressure. To alleviate this bottleneck, hierarchical KV cache offloading and scheduling are imperative. Existing scheduling policies typically rely on request histories or agent importance, without fully distinguishing task-specific return behavior within the same tool. This semantic ambiguity can cause scheduling mismatches that underutilize hierarchical storage resources and degrade response latency. To bridge this gap, we propose Model as Inference Advisor (MIA), a novel inference system that leverages the LLM itself to expose performance-relevant task semantics as structured hints, enabling highly efficient hierarchical cache management. Experiments on Terminal-Bench demonstrate that MIA reduces mean time to first token (TTFT) and mean job completion time (JCT) by up to 28.7% and 9.1%, respectively, relative to the adapted vLLM baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.