The Imperfect Mirror: Non-intrusive Model Auditing via Information Bottleneck Deterioration
Abstract
Existing non-intrusive model auditing methods mainly rely on fingerprinting, which identifies stolen models through behavioral similarity preserved from the victim model. However, this similarity is fragile under practical stealing and evasion, such as cross-architecture extraction and adversarial training. Given these limitations, we propose a novel non-intrusive framework that shifts auditing from preserved similarity to structured information deterioration. Our key insight is that independently trained models tend to exhibit a characteristic information-bottleneck balance, reflecting how they compress input information while preserving label-relevant information. In contrast, model stealing can disrupt this learned balance, leading the stolen model to retain more input information while preserving less label-relevant information than the victim, which provides a principled auditing signal. We formalize this insight into a two-stage auditing procedure that first screens for this directional deterioration and then statistically assesses whether the suspect's information-bottleneck behavior is consistent with that of independently trained models, yielding a quantitative auditing score. Experiments demonstrate effective auditing with low false-positive rates and improved robustness over the evaluated baselines under cross-architecture stealing and adversarial-training evasion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.