acceptodds
Under review as a conference paper at ICLR 2027

IndexServe: Indexer as a Service for Efficient Sparse LLM Serving

Abstract

Sparse attention reduces long-context decoding costs, but the decoder-resident index still grows with context length, limiting batch sizes even after key–value (KV) cache offloading. We present INDEXSERVE, a shared indexer service that moves index storage and token selection out of decoders while preserving the model’s selection rule. The service retains historical indexes and exchanges only current-step inputs and selected token IDs during decoding. A graph-compatible runtime and direct device-to-device transfers over a superpod fabric support per- layer remote selection, while PagedIndex and length-aware batching pool index memory and selection work across decoders. INDEXSERVE is independent of KV offloading but complements it to enable larger decode batches. Evaluated with KV offloading on DeepSeek-V3.2 and GLM-5.1, INDEXSERVE achieves 2.46–8.32× the peak generation throughput of ECHO, the strongest evaluated baseline, after accounting for service accelerators, at the cost of higher inter-token latency

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.