acceptodds
Under review as a conference paper at ICLR 2027

Skeleton Keys: Estimating Randomized Model Selection Distributions With Limited Queries

Abstract

Advances in adversarial machine learning have demonstrated that randomly selecting between defenses at inference time can further boost robustness to evasion attacks, in the white-box setting. However, it is an open question as to how an attacker could learn the model selection probability distribution in this setting. In this paper, we propose a new methodology, skeleton keys, for estimating model selection probability distributions. Skeleton keys are special inputs that leak probability distribution information to the attacker. Our contributions are as follows: First, we develop the first ever skeleton key generation algorithm and define the security properties necessary for an input to serve as a skeleton key. Second, based on skeleton key queries, we derive a likelihood estimator applicable to any defense distribution. We calculate the confidence bounds on this estimator in terms of the number of queries. Third and lastly, we empirically show the validity of our approach using state-of-the-art randomized defense ensembles. Our experiments are conducted using TRADES adversarially trained models, FAT adversarially trained models, MAMBA models and Spiking Neural Networks on CIFAR-10. Code for all our experiments is available anonymously https://anonymous.4open.science/r/proj-skeleton-key-06CE.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.