When Secrets Reveal Themselves: Vulnerabilities of Covariant Transformations in Private LLM Inference and a Low-rank Remedy
Abstract
Covariant transformations provide an efficient approach to privacy-preserving large language model (LLM) inference by applying secret transformations to private inputs and intermediate representations for privacy protection, while correspondingly transforming the model weights to preserve the original plaintext computation. However, preserving the plaintext computation through gated feed-forward networks (FFNs) and residual connections imposes common structural constraints on existing covariant transformation schemes. With the proposed CrossWeight, we show that these constraints make it possible to derive an invariant by composing transformed FFN weights, enabling recovery of the secret transformations and then the reconstruction of the private inputs. Across multiple covariant transformation schemes and model architectures, CrossWeight achieves text-token recovery rates approaching 100%, revealing a critical privacy vulnerability. To mitigate this, we further explore an irreversible transformation mechanism based on trainable low-rank transformations, together with an attack-aware objective to guarantee privacy and teacher–student distillation to preserve model utility. Extensive experiments demonstrate that the proposed method substantially reduces the effectiveness of CrossWeight and existing attacks while achieving a favorable privacy–utility trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.