acceptodds
Under review as a conference paper at ICLR 2027

Just Read the Caption: A Black-Box Privacy Attack on Knowledge Inference

Abstract

In this work, we demonstrate that vision classifiers leak private and potentially sensitive information about their semantic domain and class labels even under restricted black-box access. We show that softmax confidence acts as a natural in-distribution discriminator that an attacker can exploit to isolate relevant images from noisy web collections, using their captions to infer the meaning of otherwise opaque predictions. We introduce KEEP, a black-box knowledge inference attack that requires only the top-1 class identifier and its confidence, without prior knowledge of the victim's domain, training data, or class names. KEEP probes the victim with general-purpose image-caption pairs and uses a language model to infer its domain from the highest-confidence evidence. This domain then guides the construction of targeted probing pools from which to recover individual class labels. Across fourteen classifiers spanning four architectures and diverse domains, KEEP recovers exact class names with high accuracy, including copyrighted characters, trademarked brands, and personal identities. These results reveal a largely unexplored source of information leakage, with significant implications for the privacy and confidentiality of models trained on sensitive or proprietary data. Our code is available at https://anonymous.4open.science/r/black-box-knowledge-inference-C195/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.