Same Audio, Different Boundaries: Inferring Context, Drawing the Line, Guarding Privacy
Abstract
udio-capable models are increasingly used in applications such as personal assistants, raising privacy concerns that extend beyond conventional personally identifiable information (PII) to whether information is disclosed appropriately. This contextual dimension remains underexplored in audio scenarios. We define Contextual Audio Privacy (CAP) from an information flow perspective, where permissible disclosure is determined by the interaction context. To study CAP systematically, we construct CAP-data, a dataset of 9,000 instances spanning three settings with progressively richer contextual constraints. Empirical results show that privacy violations are widespread across the 13 evaluated audio-capable models. Privacy behavior varies substantially with requester authorization and intended use, while downstream task has a much smaller effect. Motivated by these findings, we propose FlowInspect, a runtime enforcement controller that separates privacy decisions from task execution. FlowInspect parses the relevant interaction context into audio speaker, requester role, and intended use, and maps it through a deterministic policy engine to an explicit disclosure level. Experiments show that FlowInspect reduces average privacy violation rate (PVR) from 56.82% to 31.27% across four target models while retaining 78.18% of Vanilla task successes. A 0.6B context parser retains 93.7% of the 8B parser's exact match accuracy with substantially lower memory cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.