Task-Aware Representation Release with Local Differential Privacy for Multimodal VLM Inference
Abstract
Edge–cloud collaborative inference reduces on-device computation by sending locally encoded image and text tokens to a cloud-hosted vision-language model (VLM) backbone, but these uploads may reveal private input to an untrusted server. We seek to protect private images and contextual text jointly with local differential privacy (LDP), which randomizes the upload on the client without trusting the server. However, this perturbation can damage evidence needed for answering, and directly adding noise to high-dimensional representations can introduce substantial distortion. Preserving utility under LDP therefore presents three coupled challenges: (i) reducing the dimensionality of the representation to be perturbed while retaining task-relevant evidence; (ii) configuring perturbations for differently scaled visual and textual representations; and (iii) keeping perturbed representations usable by an unchanged cloud backbone. We address these challenges through a task-trained client-side compression–perturbation–decoding framework. It combines public-question-guided attention, a joint ellipsoidal Gaussian mechanism with calibrated modality scales, and noise-aware reconstruction of the frozen backbone's native token interface. The complete image–text upload satisfies -LDP conditional on the public question and declared token structure. Only the client-side release modules are trained; the pretrained encoders and cloud backbone remain frozen. With InternVL3.5-8B, our framework improves ScienceQA accuracy by 12.2–23.5 percentage points and keyword exact match on a single-image WebQA subset by 13.8–39.2 points, compared with direct LDP perturbation of uncompressed representations at the three primary matched privacy budgets. Additional experiments show gains with a different VLM and reduced content recovery relative to unprotected uploads under the evaluated inversion attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.