acceptodds
Under review as a conference paper at ICLR 2027

UAQB: Utility-Aware Quantization Bottleneck for Private Split VLM Inference

Abstract

Split inference allows resource-constrained devices to offload vision-language model (VLM) computation to a server by transmitting intermediate visual tokens instead of raw images, thereby avoiding direct exposure of private image content. However, these visual tokens remain susceptible to reconstruction attacks because substantial spatial and semantic information is retained in them. Existing defenses mitigate leakage largely at the expense of downstream task performance. To address this challenge, we propose the Utility-Aware Quantization Bottleneck (UAQB), a plug-in defense deployed on-device prior to token transmission. UAQB applies adaptive quantization to visual tokens to restrict image information flowing through the bottleneck, coupled with utility adaptation that restores compatibility of the resulting information-limited representation with the downstream VLM. We derive a tractable upper bound on the information bottleneck objective, enabling joint optimization of privacy and utility. Extensive evaluations across multiple utility benchmarks and reconstruction metrics demonstrate that UAQB achieves a more favorable utility-privacy trade-off than existing defenses, with particularly high utility retention at stronger protection levels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.