acceptodds
Under review as a conference paper at ICLR 2027

Anamorph: A Secure and Efficient TEE-Shielded Obfuscation Scheme for On-Device LLM Inference

Abstract

Deploying large language models on untrusted devices poses a risk of model weight leakage, and Trusted Execution Environment (TEE)-based model obfuscation provides a practical solution for protecting model weights. The state-of-the-art lightweight defense based on dense orthogonal matrices can effectively resist cosine-similarity-based analysis attacks while reducing TEE-side computation. However, we identify a key vulnerability in existing lightweight obfuscation schemes: statistical similarity between private fine-tuned models and public pretrained models. Furthermore, we find that the problems of estimating secret transformation matrices in multiple existing obfuscation schemes can be reduced to orthogonal Procrustes problems. Based on these observations, we propose ORTHOTRACE, a novel attack that exploits the correspondence between public pretrained weights and exposed obfuscated weights to estimate secret transformation matrices, thereby compromising multiple obfuscation mechanisms, including the state-of-the-art lightweight defense. To mitigate this vulnerability, we propose ANAMORPH, a novel obfuscation scheme that combines parameter-dimension expansion with the injection of random masks satisfying null-space constraints to perturb the orthogonal estimation relationships exploited by ORTHOTRACE. These masks vanish automatically during authorized inference, reducing the need to enter the TEE for layerwise mask correction and maintaining lightweight TEE computation while ensuring model security. We evaluate ORTHOTRACE and ANAMORPH on four representative models and eight datasets. Experimental results show that ORTHOTRACE compromises multiple existing lightweight obfuscation schemes, recovering models with task performance close to the no-protection baseline, with the task recovery rate exceeding % for each scheme. ANAMORPH demonstrates effective defense against existing similarity-based analysis attacks and ORTHOTRACE while preserving authorized-inference utility, and achieves up to a speedup in time to first token (TTFT) compared with an effective defense scheme that requires layerwise correction on real SGX hardware.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.