acceptodds
Under review as a conference paper at ICLR 2027

PrivacyNav: Controlling Visual Representation Disclosure in Vision-Language Navigation

Abstract

Vision-Language Navigation (VLN) systems rely on semantically rich visual representations for scene understanding and action generation. However, in modular, cross-device, or remote inference, these representations are often transmitted beyond the trusted perception boundary, where an honest-but-curious recipient could recover policy-defined sensitive information. To address this issue, we propose PrivacyNav, a state-adaptive visual disclosure framework that controls which visual evidence is disclosed and how much is disclosed at each navigation state, while keeping the underlying VLN policy frozen. PrivacyNav comprises two modules. Privacy-Aware Evidence Estimator predicts token-level disclosure safety together with state-level disclosure risk, yielding localized and contextual signals. Adaptive Visual Disclosure Controller ranks visual tokens by jointly considering disclosure safety and host-native navigation evidence, and maps state-level risk to a state-specific token retention ratio. We conduct extensive experiments across four heterogeneous VLN architectures, complemented by module ablations, operating-point analyses, and real-world deployment. Compared with disclosing unmodified representations, PrivacyNav reduces macro sensitive-concept recovery by 3.23–10.38 percentage points, with navigation success rate losses of 0.81–5.11 percentage points. These results demonstrate consistent reductions in sensitive semantic recoverability under the evaluated protocols, with host-dependent disclosure–utility trade-offs rather than a formal privacy guarantee.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.