SparseEngine: Sparse-First Inference Engine
Abstract
Long-context LLM-based agents accumulate growing interaction histories, placing substantial pressure on KV cache memory and attention computation. Consequently, recent works have adopted sparse attention to reduce computational and memory overheads. However, due to the heterogeneous nature of sparse methods, they cannot be directly applied to existing inference engines. *Sparse inference engines* have therefore been proposed, but existing serving abstractions tie method integration to particular cache layouts or computational workflows, limiting support for heterogeneous sparse methods. We present **SparseEngine**, a ground-up sparse-first inference engine built around a shared lifecycle contract. The contract lets each method control its KV representation and computational workflow while coordinating state transitions with shared serving infrastructure. SparseEngine supports **15** sparse methods across four major categories, demonstrating its generality. Furthermore, built upon its lifecycle contract, it enables higher-level cross-request state management: we propose **Chain Cache** to allow KV eviction methods to resume from retained history, and Controllable **Prefix-Cache Pruning** to let applications prune KV from selected history regions while preserving logical prefix matching. While preserving the quality of sparse methods, SparseEngine achieves over **10×** higher throughput with KV eviction methods and over **2.5×** decode throughput speedup under identical concurrency, compared to vLLM. It also achieves over **2×** end-to-end speedup on agent benchmarks. The code is available at [https://anonymous.4open.science/r/SparseEngine](https://anonymous.4open.science/r/SparseEngine).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.