Can Active Rival Interactions Expose Adversarial Vulnerabilities in Deep Networks?
Abstract
Adversarial attacks exploit the decision geometry of deep neural networks to expose vulnerabilities under small input perturbations. However, these attacks primarily optimize top-1 misclassification losses or manipulate ranked outputs, leaving dynamic pairwise interactions among non-source rival classes closest to overtaking the source class underexplored, which may, in turn, limit attack effectiveness. In particular, probing a direction associated with one rival class can reveal changes in another rival class’s normalized margin gradient. This reveals a potential source of attack potency: pairwise margin-field responses provide geometric information beyond individual rival gradients at the current iterate and can expose adversarial directions missed by independent classwise optimization or single-margin ascent. We therefore introduce Rival Interaction Optimization for Holistic Attack (RoHoAttack), a geometry-aware white-box attack that exploits active top-K non-source rival-class selection, reciprocal probing, residual refinement, and projected lookahead, thereby retaining the highest-margin evaluated ℓp-feasible candidate while refreshing the active rival classes after each update. Across natural-vision and medical-imaging benchmarks, RoHoAttack achieves higher attack success rates and reduces robust accuracy by up to ≈ 12.5 points over leading baselines, thereby demonstrating its attack effectiveness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.