The Mean Is Not Enough: Defeating Query-Based Attacks with Learned Disagreement
Abstract
Many defenses against query-based adversarial example (AE) attacks are vulnerable to Expectation over Transformation (EOT): by averaging repeated responses, an attacker can estimate the defense's mean response and optimize directly against it. But recovering the mean is useful only when AEs crafted against the mean transfer reliably across the predictors. We present DISSENT, which suppresses this transferability by learning a distribution of predictors that agree on clean inputs but disagree along adversarial directions. Its learning objective discourages AEs crafted against one sampled predictor from transferring to another, while preserving agreement on benign inputs. At inference, each query is answered by one randomly selected predictor in a single forward pass, avoiding the computational cost of evaluating multiple predictors for voting or averaging. Our theoretical analysis explains how learned disagreement can sustain robustness even when the attacker knows the mean. We evaluate DISSENT on CIFAR-10 and ImageNet with different architectures. On CIFAR-10 with ResNet-50, for example, DISSENT retains 59.5% robust accuracy under EOT, compared with 23.1% for the strongest evaluated baseline. Further evaluations against adaptive attacks, including those crafted on method-aware surrogates, demonstrate its robustness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.