Beyond Static Masks: Attacking and Defending TEE-Assisted LLM Inference
Abstract
TEE–GPU collaborative inference enables efficient deployment of proprietary large language models (LLMs) by combining trusted CPU execution with untrusted GPU acceleration. Recent approaches, such as SLIM, reduce TEE–GPU interaction through transformed-domain propagation across decoder blocks. However, we reveal that this propagation mechanism introduces a transformation identifiability vulnerability under multi-query inference. By exploiting reusable linear relationships exposed during authorization, an adversary can recover equivalent secret transformations and compromise the protection without accessing the original secrets. To address this issue, we propose GateFlow, a secure authorization propagation framework that redesigns the initial decoder execution while preserving efficient transformed-domain propagation. GateFlow prevents recoverable intermediate representations during authorization and enables subsequent decoder blocks to inherit the propagation mechanism without additional TEE involvement. Extensive experiments validate the effectiveness of the proposed attack and demonstrate that GateFlow effectively mitigates transformation recovery attacks with practical inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.