Beyond Independence: Modeling Correlated Adversarial Transfer In Model Ensembles
Abstract
Ensembles are widely believed to resist black-box transfer attacks: if adversarial examples transfer imperfectly, compromising several members at once should be rare. This argument silently assumes that transfer events across members are independent. We test that assumption directly, treating transfer against an ensemble as a vector of correlated binary events and measuring per-model transfer rates, pairwise correlations, and the overdispersion of the fooled count for six-model ensembles on CIFAR-10, GTSRB and MNIST under PGD transfer from a ResNet surrogate. In the regime where the ensemble argument is actually invoked, at small budgets where per-model transfer is a few percent and majority compromise is meant to be rare, independence is rejected wherever the data are well powered, including by an exact permutation test (): the observed rate of compromising a majority exceeds the independence prediction by more than an order of magnitude ( on CIFAR-10, on GTSRB, on MNIST at the smallest budget with events). As the budget rises and transfer saturates the gap closes and eventually reverses, exactly as it must; the danger zone for a defender is the low-budget corner, and that is where independence fails hardest. The failure is correctable with a single parameter: an exchangeable beta-binomial calibrated on the measured overdispersion matches the observed risk in every well-powered configuration, where independence fails in all of them, and, fit at one ensemble size, predicts majority-vote risk at other sizes to within a percentage point. The dependence is not an artifact of our training pipeline: it survives replacing distilled students with independently trained ones and the teacher surrogate with an unrelated model, and it persists under three transfer-optimized attacks. It also survives the standard fix: adversarial training lowers how often members are fooled but not how correlated their failures are, so the residual risk must still be read through the dependence structure. Ensemble robustness is governed by the dependence structure of failures, not by per-model transfer rates alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.