VisDelta: Model Inversion from Released LoRA Adapters for Vision-Language Models
Abstract
Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning (PEFT) method for vision-language models (VLMs) that updates only a small fraction of their parameters. When trained on private images, LoRA adapters retain visual information from them. Adapters are commonly shared on public model platforms, creating a serious risk of model inversion. In this paper, we propose VisDelta, a model inversion attack against released VLM LoRA adapters. It compares the base and LoRA-adapted VLMs on the same image and query and uses their response difference to guide image reconstruction. Firstly, VisDelta samples latent codes at random and generates an initial image from each using a public image generator. It ranks these images by an adapter-aware score, retains the codes that produced the top-ranked images, and refines each code through adapter-aware differential guidance to reconstruct private images associated with the target. The guidance combines target agreement, separation from hard-negative conditions, and the adapter-induced gain in target separation, while keeping each code close to its initial value. Refinement uses random image transformations; final selection balances average and worst-view losses over fresh transformations to choose a reconstruction with stable target responses. Experiments across four VLMs and three benchmark datasets show that VisDelta recovers target-related visual content from LoRA adapters trained on private images. On CelebA with Qwen2.5-VL-7B, it achieves 26.38% Acc@1 and 52.47% Acc@5, exceeding the strongest baseline by 3.81 and 2.29 percentage points, respectively. These results show that released VLM LoRA adapters can expose private visual information.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.