Taurus: Accelerating Private LLM Inference via Homomorphic Speculative Decoding
Abstract
Cloud-hosted large language model (LLM) inference raises privacy concerns, since sensitive user inputs are exposed to the service provider. Fully homomorphic encryption (FHE) offers a promising solution by enabling inference directly over encrypted inputs. However, its substantial computational overhead remains a major obstacle to practical deployment. This overhead is particularly severe in autoregressive decoding, which generates only one token at a time, while FHE is more efficient when many values are packed and processed together. Speculative decoding instead processes multiple tokens in one model evaluation pass, creating an opportunity to better utilize FHE's packing ability. However, its speculation-specific operations, including sampling, verification, and state updates, are expensive to realize under FHE. In this paper, we propose Taurus, the first framework for fully homomorphic speculative decoding. Taurus packs many speculative candidate tokens together to share expensive FHE computation across multiple token states. To be more efficient and accurate, Taurus develop a homomorphic sampling protocol that avoids scanning the entire vocabulary with encrypted comparisons. Taurus also proposes a private speculative verification protocol to avoid encrypted division or clipping, and performs state updates obliviously without revealing which candidate tokens are accepted. Together, these protocols enable efficient speculative decoding entirely on the server, reducing per-token latency while preserving the privacy for both client and server. We implement Taurus and conduct experiments on six open-weight models with 1.7B to 8B parameters. On eight NVIDIA RTX 4090 GPUs, Taurus achieves a 1.80× to 3.02× speedup over CacheMir in end-to-end decoder latency per committed token.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.