Mitigating Automated Geolocation via Semantic-aware Typographic Context
Abstract
Large Vision-Language Models (LVLMs) have demonstrated a strong ability to infer geolocation directly from images, raising privacy concerns as increasingly accessible LVLMs lower the barrier for large-scale location inference from publicly available images. Existing approaches for mitigating such risks primarily rely on adversarial image perturbations, which require substantial visual distortions and consequently degrade image quality. This limitation motivates a unexplored research question: Can unintended geolocation inference be mitigated through contextual information beyond the image itself? To investigate this question, we explore which textual semantics are effective in influencing LVLM-based geolocation reasoning and design a two-stage, semantics-aware typographical attack that generates contextually misleading text, termed GeoSTA. As a platform-side intervention, GeoSTA generates a protected image derivative for public or automated access by introducing semantic typographic context outside the original image region. Extensive experiments across three datasets and five commercial LVLMs demonstrate that GeoSTA significantly disrupts geolocation inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.