Traces of a Prompt: Uncovering and Mitigating Privacy Risks of Exposed States in TEE-Partitioned LLM Inference
Abstract
Cloud-based LLM services must protect sensitive prompts while maintaining efficient inference. To balance these goals, partitioned inference keeps selected computations inside a trusted execution environment (TEE) while offloading others to an untrusted GPU. However, the intermediate states exposed by this partitioning may reveal private inputs. We introduce Prompt-State Matching Attack, which constructs direct, invariant, and joint-state fingerprints for vocabulary-based prompt reconstruction. Given observations of selected states and access to chosen-input probing, the attack achieves near-perfect token recovery through exhaustive vocabulary search in controlled FP32 experiments across four LLMs. We experimentally evaluate token-aligned transformations and analyze how relationships among exposed states support reconstruction. To mitigate this leakage, we propose StateVeil, which protects outsourced linear computation using Dual-LPN-based sparse encoding. We provide a conditional reduction for transcript privacy, complemented by concrete attack-cost estimates. Even with the true prefix supplied at each position, the evaluated Euclidean direct matcher achieves token recovery from StateVeil's masked outputs across four models and four protected projections. StateVeil achieves end-to-end latency comparable to the evaluated partitioned-inference baseline, while model quality remains comparable to the original models in most evaluated model–task settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.