acceptodds
Under review as a conference paper at ICLR 2027

Differentially Private Multimodal In-Context Learning

Abstract

Vision-language models (VLMs) are increasingly deployed on sensitive data such as medical scans and financial documents, where in-context learning offers customization without retraining. Yet existing differentially private in-context learning methods are limited to few-shot, text-only settings because privacy cost scales with the number of demonstrations, and a single image can consume hundreds of tokens. We introduce Differentially Private Multimodal Task Vectors (DP-MTV), the first framework for many-shot multimodal in-context learning with differential privacy. Rather than protecting individual demonstrations, DP-MTV aggregates hundreds of private image-text examples into a single steering vector and privatizes the aggregate. Privacy cost is paid once at construction; the released vector then answers unlimited queries at zero additional cost. Across eight benchmarks and three VLM architectures, DP-MTV matches or exceeds non-private performance on the majority of tasks at epsilon=2, and on classification consistently outperforms the non-private baseline. Beyond this public-data setting, a fully-private variant that uses no auxiliary data for head selection matches it at epsilon=5, (for example, 96% on Flowers-102 with Qwen-VL versus 77% for non-private MTV)with a per-layer selection scheme extending competitive accuracy to tighter budgets. A novel membership inference attack validates these guarantees: 100% success without protection drops to the 50% random-guess baseline under DP-MTV.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.