acceptodds
Under review as a conference paper at ICLR 2027

You Can't Read My Mind: Non-interactive Secure Autoregressive LLM Inference

Abstract

Large Language Models (LLMs) excel at autoregressive generation but typically require server-side deployment, raising the risk of sensitive-data leakage. Fully Homomorphic Encryption (FHE) enables secure LLM inference directly over ciphertexts. However, autoregressive inference under FHE remains underexplored and faces challenges in both efficiency and generation quality. Token-by-token generation requires expensive vocabulary-wide operations, while approximate token selection over vocabulary introduces semantic errors and unintended mixtures of token embeddings. These operations recur at every decoding step, making their computational costs and approximation errors a bottleneck to a practical inference under FHE. To address these challenges, we propose , a non-interactive framework for secure autoregressive LLM inference. We introduce a novel inference flow: the server performs reasoning over ciphertexts, while a lightweight client model generates the answer locally with the returned reasoning latents. This design keeps vocabulary-dependent operations outside FHE evaluation. Fixed-length reasoning in encrypted latent space further reduces encrypted decoding from an average of 26.7 steps to 5 on GSM8K-Aug. To accelerate the decoding stage under FHE, we introduce a layout-compatible KV-cache decoding algorithm that efficiently reuses prefill KV-cache without packing-format conversion. We implement TELEX as an end-to-end system that achieves near-lossless generation quality and substantial efficiency gains. TELEX achieves an per-step decoding speedup compared to MOAI. On a single H200 GPU, we report an amortized time of 49.9 seconds per query for decode stage and an end-to-end time of 14.8 minutes per query. TELEX is the first open-source system to support end-to-end, non-interactive autoregressive LLM inference under FHE.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.