RheoServe: CPU-GPU Collaborative Sparse Attention for Efficient LLM Decoding
Abstract
Sparse attention is a promising approach for efficient LLM decoding over large KV caches, yet managing these caches remains a bottleneck: eviction-based methods sacrifice accuracy, while retrieval-based offloading is constrained by limited PCIe bandwidth, thereby bottlenecking decoding throughput. To address this, we present RheoServe, a CPU–GPU collaborative sparse-attention framework that actively leverages host memory bandwidth via block-level cache management. RheoServe adopts a compute offload paradigm: rather than moving massive data to the GPU, it moves compute to the data, executing sparse attention directly on DDR-resident blocks and merging results via Log-Sum-Exp (LSE) reduction. Furthermore, RheoServe employs a unified cache mechanism that intelligently manages block placement, improving cache hit rate to compensate for the bandwidth gap. Extensive evaluations demonstrate that RheoServe maintains reasoning accuracy and achieves and the peak throughput of matched PCIe Retrieval for Qwen3-4B and Qwen3-8B, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.