FactPatch: Persistent Visual Pseudonymization for Multi-Turn Vision-Language Models
Abstract
Multimodal large language models may repeatedly access the same sensitive fact in an image over the course of a conversation, and these accesses are not independent. We find that an early visual fact interpretation becomes part of the model-generated dialogue history and can subsequently influence how the same fact is interpreted, causing identical images to evolve toward different and persistent interpretations along different conversational paths. This suggests that multi-turn visual privacy is not merely about changing a single response, but about maintaining a protected fact interpretation under continuously changing conversational evidence. Based on this observation, we propose FactPatch, which uses a small, localized visual perturbation to map a true sensitive fact to a semantically consistent substitute and leverages persistent visual evidence to stabilize this interpretation across diverse conversational paths and factual disturbances. When the model drifts back toward the true fact, the same visual evidence helps restore the protected interpretation. Experiments show that FactPatch more reliably preserves the substitute fact on unseen conversational trajectories and held-out models, substantially reduces the re-exposure of true sensitive information over multi-turn interactions, and largely preserves the model’s ability to understand non-sensitive visual content.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.