Defending Wearable VLMs Against Private Attribute Inference
Abstract
Split vision-language model (VLM) assistants may transmit visual tokens from a trusted edge device to a downstream language model. These tokens support the user's requested task but can also encode private attributes of the wearer or nearby people. A source model that avoids stating those attributes may still expose them through its transmitted tokens to a newly trained attacker. We study this privacy–utility tension using a paired benchmark of 3,221 image-question records from Reddit posts and converted VQA datasets. Each record has a utility question and available evidence-supported labels for location, income, sex/gender, or interests. The benchmark is a single-image proxy for wearable assistance, not smart-glasses footage or continuous first-person video. We propose Token-Guided Attribute Privacy (TGAP), a lightweight residual adapter that transforms post-resampler visual tokens inside the trusted boundary before transmission while keeping the source VLM frozen. Its objective combines utility preservation and identity regularization with source-model privacy suppression and an auxiliary private-attribute decoder that acts on token representations. On MiniCPM-V, TGAP reduces source-model privacy accuracy from 56.7% to 7.4%, a 49.3% reduction, while relaxed utility changes from 76.6% to 74.4%. Separately, fresh attacker bridges and heads trained directly on filtered tokens show lower average attribute recoverability across four frozen surrogate LLM families. These results support empirical mitigation at the token interface, while residual leakage and attacker-dependent outcomes remain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.